Skip to content

AI council

Four judges score every finding. None of them talk to each other.

Instead of one model deciding your risk score, AdverseMe asks four models from four providers to rate each finding independently, takes the median, and shows you every vote.

Four independent scores frozen on the same evidence

The AI council is how AdverseMe turns evidence into a risk score. 4 language models from 4 different providers each rate every finding on impact, likelihood and control weakness at temperature zero. The median becomes the score, every vote is kept, and identical evidence returns an identical result for 14 days.

Judges
4, from 4 model families
Sampling
Temperature 0, fixed seed per judge
Score
Median of the judges
Same evidence
Same score for 14 days

How it works

From evidence to a score you can reproduce

The council only ever scores findings. It never searches, and it never sees another judge's answer.

  1. 1

    Evidence is collected first, without any scoring

    The engine searches 117 sources, resolves which results are really about your subject, and drops same-name matches for a different person or company. Nothing here involves a judge.

  2. 2

    Concrete findings are extracted

    Sanctions matches, court records, adverse media, PEP positions and sanctioned associated parties are flattened into a numbered list of findings. Neutral news and directory listings are excluded, so the judges only ever see adverse signals.

  3. 3

    4 judges score each finding independently

    Each seat on the council is a different model from a different provider. Every judge receives the same finding list and returns impact, likelihood and control weakness for each item plus an overall rating. Temperature is locked to zero and each judge has its own fixed seed. If a provider fails, that seat retries once on a fallback and is otherwise counted as unavailable rather than guessed.

  4. 4

    The median wins and dissent is counted

    Per finding, the median impact, likelihood and control weakness are multiplied and normalised to 100. The screening score is the median of the judges' overall scores. Judges whose rating differs from the median are counted as dissent, and each finding is marked unanimous, majority or split.

  5. 5

    The verdict is fingerprinted

    A signature of the subject details, the current sanctions list versions and the scoring methodology version is stored with the result. Re-screen the same subject within 14 days while nothing upstream has changed and you get the same score, not a fresh roll of the dice.

What you see in the app

A findings table that opens into every judge's vote

#FindingILCWScoreRatingConsensus
1Sanctions list match55480CriticalAll agree
2Regulator enforcement notice43329Medium3/4
Expanded row
j1 (Google) · I 4 L 3 CW 3 · Medium · "Fine was paid and the licence remains active."
j2 (OpenAI) · I 4 L 3 CW 3 · Medium · "Single enforcement action, no repeat."
j3 (Meta, open weights) · I 4 L 4 CW 3 · High · "Pattern suggests weak controls."
j4 (Anthropic) · I 4 L 3 CW 3 · Medium · "Remediated, disclosed publicly."

Illustrative values. I = impact, L = likelihood, CW = control weakness, each 1 to 5. Score is the product normalised to 100.

The screening result page lists every finding with the median component scores, the normalised score and the rating. A consensus chip shows All agree, 3/4 or Split.

Click a row and the per-judge table opens underneath: judge id, provider, model, the three component scores, the rating and the judge's reasoning text. If a seat was unavailable it says so.

The headline verdict uses the same numbers. The compliance interpretation shown at the top is pulled from the judge whose overall score sits closest to the median, so the narrative matches the number.

Why it matters to a regulator

When someone asks why it was rated Medium, you have four answers

Supervisors and auditors do not accept "the AI said so". They want to know what was assessed, by what method, and whether the outcome would be the same tomorrow. The council gives you all three: a fixed finding list, a published method (impact times likelihood times control weakness, median across judges, severity-weighted overall), and a stored signature that proves the score is reproducible on unchanged evidence.

Dissent is kept on purpose. A split vote is a signal that the evidence needs a human, and the record shows you saw it. An unscored result is labelled as unscored, never silently zero.

The PDF report deliberately leaves the machinery out and states facts with a source and a date. The per-judge breakdown stays in the app for your own file. See Reports and Engine Trace.

Frequently asked questions

Why four judges instead of one model?

One model can have a bad day on one prompt. 4 independent models from 4 providers, each at temperature zero, give a median that is far harder to skew, and the dissent count tells you when the evidence is genuinely ambiguous.

What happens if a judge is unavailable?

The seat retries once on a fallback model. If that also fails, the seat is recorded as unavailable and the median is taken over the judges that answered. If no judge answers at all, the screening is reported as not scored rather than given a made-up number.

Is the score deterministic?

Yes within the evidence window. The council runs at temperature zero with fixed seeds, and a signature of the subject plus the sanctions list versions is stored with every result. A repeat screening within 14 days on unchanged evidence reuses the same verdict.

Can I see how each judge voted?

Yes. In the app every finding expands into a per-judge table with the provider, model, the three component scores, the rating and the reasoning text. The consensus chip on each row shows how many judges agreed.

See the four votes on your next screening.

Plans start at $49 a month for 50 screenings. Every plan includes the AI council, the ownership graph, Engine Trace and PDF reports.