Skip to content

Why AdverseMe scores every finding with four judges at temperature zero

One model scoring adverse media is an opinion. Four models from four families, run blind at temperature zero with fixed seeds, with the median taken and dissent counted, is a reproducible measurement. Here is why we built it that way.

  • ai-council
  • explainability
  • product

Ask a large language model whether a news article is adverse media about your customer and you will get an answer. Ask it again tomorrow and you may get a different one. Ask a different model and you will get a third. For a compliance decision that a regulator can revisit two years later, that is a problem. This post explains how AdverseMe’s AI council is built, what each design choice is for, and what it means for the analyst reading the score.

Key takeaways

  • Every finding is scored by four judges drawn from four different model families, none of which sees the others’ output.
  • Each judge runs at temperature zero with a fixed seed, so the same evidence produces the same score within a 14-day reproducibility window.
  • The reported score is the median of the four; the spread between judges is reported as dissent rather than hidden.
  • Every judge’s reasoning is written into the report, so a reviewer can disagree with the council on the evidence, not on faith.

The problem with one model

A single model scoring adverse media has three weaknesses that matter in compliance.

The first is variance. Language models sample from a probability distribution. At the default settings, two runs on identical input can differ, and a “7 out of 10” today can be a “5” tomorrow with nothing changed but the dice. A risk score that moves on its own is not a measurement.

The second is correlated error. Every model family has blind spots: names it over-associates with crime, languages it reads poorly, article structures it misreads. If one model produces the score, its blind spots are the product’s blind spots, and there is no second opinion to catch them.

The third is explainability. A score with no reasoning attached tells an MLRO nothing they can defend. “The model said 8” is not a finding.

Four judges, four families

AdverseMe scores each finding with four judges. The judges come from four different model families, deliberately, so that their errors are as uncorrelated as we can make them. A family-specific bias in one judge is unlikely to be shared by the other three.

Each judge sees the same evidence: the article or record, the subject’s identifying details and the context of the screening. None of them sees the other judges’ scores or reasoning. This is what we mean by scoring blind. There is no chain where one model summarises and another grades the summary; each judge reads the primary evidence and forms its own view.

The output from each judge is a score and a written rationale. The rationale is not decoration. It is kept, shown to the analyst and printed in the PDF report beside the score.

Temperature zero and fixed seeds

Temperature is the setting that controls how much a model samples away from its most likely output. At temperature zero the model returns its most likely completion, and with a fixed seed the sampling path is pinned. AdverseMe runs every judge at temperature zero with fixed seeds.

The effect is determinism: identical evidence returns an identical score. We hold that within a 14-day reproducibility window, which is the period over which we commit that a re-run of the same screening against the same evidence produces the same council result. The window exists because model providers update their models; the commitment is honest about that rather than pretending that a score is reproducible forever.

Why does this matter? Because a compliance decision has to be defensible after the fact. If a regulator asks why a finding was scored a 4 and cleared, the answer “run it again and you will get a 4, here is the reasoning” is a defence. “It was a 4 that day” is not.

The median, not the average

With four scores in hand there is a choice about how to combine them. AdverseMe takes the median.

The average is sensitive to a single outlier. If three judges score a finding 2 and one scores it 9, the average is a misleading 3.75 that represents nobody’s view. The median is 2, which is what three of the four judges concluded, and the outlier is preserved as dissent rather than blended away.

The median is also easier to explain. “The middle score of four independent judges” is a sentence an analyst can say to an auditor.

Dissent is a signal

When judges disagree, that disagreement is information. A finding that scores 2, 2, 3, 3 is a confident low. A finding that scores 1, 2, 8, 9 has the same rough middle and is nothing like as clear.

AdverseMe counts dissent and reports it with the score. A high-dissent finding deserves human review regardless of where the median lands, because a split council usually means the evidence is ambiguous: the article may concern a different person of the same name, the allegation may be old or withdrawn, or the source may be unreliable. Those are exactly the cases where an analyst’s judgement adds the most.

A tool that hid the spread and showed only the middle would be quieter, and worse.

What the analyst sees

On each finding the analyst sees the median score, the dissent, and four rationales side by side. If a judge dismissed the article because it concerns a namesake in a different country, that reasoning is visible. If another judge weighted a conviction heavily, that is visible too. The analyst can agree with the council, override it, or send it for a second opinion, and the record shows what the council said and what the human decided.

The same trail continues into the Engine Trace, which records every source searched and every candidate kept or dropped before the judges ever saw it, and into the report, where each judge’s vote and reasoning is printed for the file.

What this is not

It is not a claim that four models are always right. They are not. The design goal is narrower: to make the score reproducible, to make single-model error less likely to pass unnoticed, and to make the reasoning inspectable so that a human can catch what the council missed. Those are properties a compliance team can build a process on. Accuracy is something the team verifies case by case, and the council is built to make that verification possible.

It is also not a black box with a confidence percentage on top. Every number on the screen can be traced to a rationale and to a source.

Where it fits

The council scores adverse media and other findings from all 117 sources AdverseMe screens; the coverage page lists them. It runs on every plan, including the Starter tier, and it runs the same way for a single screening as for a batch upload. The AI council feature page shows the interface.

If you want to see four rationales side by side on a real finding, get started and screen a name you already know the answer for.

Screen the next name with the working attached.

Plans start at $49 a month for 50 screenings. Every plan includes the AI council, the ownership graph, Engine Trace and PDF reports.