elmerdata.ai blog

My blog

9 Judges, 2 Votes

Apple tested a panel of 9 frontier models from 7 model families and found the panel carried about 2 independent votes' worth of information. The obvious fix for an AI that grades its own work is a second AI. The obvious fix has a problem.


Cicero addressing the Roman Senate as Catiline sits alone on an empty bench
Cesare Maccari, Cicero Denounces Catiline, 1882 to 1888. Fresco, Palazzo Madama, Rome, now the seat of the Italian Senate. Public domain. Maccari gives the chamber a great many senators and exactly 1 opinion, which the empty bench around Catiline makes unmistakable.


Nine AI judges should be worth something close to 9 opinions. Apple tested a panel of 9 frontier models drawn from 7 model families and found something closer to 2.

The paper is Guneet Kohli's Nine Judges, Two Effective Votes, and the numbers deserve stating plainly. Roughly three quarters of the panel's nominal independence disappears because the models make the same mistakes on the same items. Accuracy falls 8 to 22 percentage points short of what independent voting would deliver. The best single judge matches or outperforms the whole panel across every condition tested. Adding judges does not help, and neither does smarter aggregation: established methods close at most 11% of the gap even when handed the correct answers.

The paper's conclusion is the sentence to keep. The bottleneck is correlated judges, not the aggregation algorithm, which means scaling up panels cannot substitute for genuinely independent evaluation.

The obvious solution

I have spent 2 posts on the problem that is supposed to solve. AI evaluators favor their own work, by margins large enough to reorder a shortlist, and the favoritism concentrates where the evaluator is wrong, because that is where its judgment stops tracking quality and starts tracking familiarity. Both posts ended in the same place. If a model cannot grade its own work, get a second opinion.

Pombal, Rei and Martins complicated that. Self-preference extends past the model to its relatives, with same-family judges showing elevated favoritism, which leaves a harder question their paper does not settle: what makes 2 models related? Shared training data, a distillation history and common post-training methods all link systems carrying different names.

Apple's result pushes further, because 7 model families is not a subtle test of diversity. So the rule most institutions are reaching for, roughly use a different model or at worst a different vendor, does not do what it appears to do.

Counting ballots

Apple's framework is effective sample size, borrowed from survey statistics, and the intuition survives without the mathematics. Picture a committee of 9. If each member reasons independently, 9 votes carry far more information than 1. Now suppose 7 of them trained under the same supervisor, read the same brief and share the same blind spots. You still count 9 ballots. You no longer have 9 judgments.

The financial version is more familiar. A portfolio of 9 securities is not diversified if all 9 respond to the same shock in the same direction, which is why investors stopped counting holdings and started measuring correlation. AI assurance has not made that move, and the architecture diagrams show it. A system with a generator, an evaluator, a second evaluator and a safety monitor appears to contain 4 controls. The meaningful question is not how many boxes are on the diagram but how correlated their failures are, because 4 boxes that fail under the same conditions are 1 failure mode drawn 4 times.

We may be counting models when we should be counting independent mistakes.

The auditors already learned this

Which brings me to a profession that solved a version of this in 2003 without asking anyone to try harder.

The SEC's auditor independence rules turn on a structural question rather than a behavioral one, and the language is direct: they prohibit an accountant from auditing the bookkeeping work performed by his or her own firm. Sarbanes-Oxley Title II went on to bar an issuer's auditor from simultaneously supplying 9 categories of non-audit service, among them financial information systems design, appraisal and valuation, internal audit outsourcing and legal services.

Notice what those rules do not say. Nobody instructs the auditor to set aside the fact that it designed the system it is now examining, or asks for extra objectivity, or adds a second partner to review the first. The relationship itself is disqualifying, and the profession arrived there by discovering that good faith is not a control.

Banking then applied the same logic to models specifically, in 2011. The Federal Reserve and OCC's supervisory guidance on model risk management reads, 15 years on, as though drafted for this argument. Validation involves a degree of independence from model development and use, it says, and is generally done by staff who are not responsible for either and who have no stake in whether a model is found valid. Its organizing idea is effective challenge, critical analysis by objective, informed parties who can identify a model's limitations, and it names what effective challenge rests on: incentives, competence and influence. Not one of those 3 is a property of the evaluator's reasoning. All 3 are properties of where the evaluator sits.

Set that against the mitigations in the AI literature, all of which are requests that the evaluator behave better. Instruct it to disregard the source. Make it reason step by step. Average several of them. Each helps, measurably. None eliminates the effect, and the field keeps reporting that as partial success.

Independence is not an instruction. It is an architecture.

What AI governance has, and what it lacks

Plenty of people in AI policy have thought about independent evaluation, and the documents are better than the debate suggests. NTIA's AI Accountability Policy Report of March 2024 says self-assessments are unlikely to be sufficient, that independent evaluations provide essential checks on management's own assessments, and that allowing developers to certify their own software is a clear conflict of interest. NIST's AI Risk Management Framework asks at Measure 1.3 that internal experts who did not serve as front-line developers, and independent assessors where appropriate, be involved in regular assessments. In July 2026 NIST launched its Artificial Intelligence Technology Evaluation program, which puts models in a sequestered testbed against blind data to mitigate train and test contamination.

The binding law is thinner than its reputation. The EU AI Act routes high-risk systems under points 2 to 8 of Annex III through an internal control procedure that, in the text's own words, does not provide for the involvement of a notified body. Article 55, the obligation attaching to general-purpose models with systemic risk, requires the provider to perform the model evaluation and document the adversarial testing. The provider performs it. The strongest external-evaluation requirement in the landscape is not in the regulation at all: it sits in the voluntary General-Purpose AI Code of Practice, which asks signatories to have independent external evaluators conduct model evaluations, with an exemption for anyone who cannot appoint a qualified one. The opt-in code asks for more independence than the binding law does.

All of it answers a single question: who should evaluate. The research above raises a second that none of these frameworks reaches. An evaluator can be organizationally independent and statistically correlated with the system it evaluates. Different company, different contract, different incentives, and the same errors on the same items.

So what is the unit of AI independence? Model instance is plainly insufficient. Model family may be insufficient. Vendor diversity may be insufficient. Training provenance probably matters, and so may distillation, shared synthetic data and common post-training methods, several of which the purchasing organization has no way to discover. Accounting knows how to prohibit certain relationships. AI does not yet know which relationships matter.

Where the average misleads

A complication arrived 2 months after Apple's paper, and it cuts both ways.

Yang Shu's Blind to the Pivotal Vote takes the nine-judges result as its starting point and argues that aggregate independence metrics miss where independence actually pays. Adding an external verification signal, such as executing a test suite, barely moves the effective vote count at all, a change of 0.04 in the wrong direction. Yet the accuracy it buys is not distributed evenly. The entire gain concentrates on the pivotal cases where the panel is genuinely split, and there it is worth 10.4 to 23.3 percentage points. Everywhere else it is exactly zero.

Which does soften the headline number. If effective votes can stay flat and accuracy still move sharply, then 9-becomes-2 is not the whole story either. What survives is the finding underneath it, and it is the same lesson the last post reached from the other side. Self-preference is smallest where a model is right and largest where it is wrong. The value of an independent signal is zero where the panel agrees and enormous where it does not. Both say the average is the wrong statistic.

The second half of Shu's result is the more useful one. The signal that broke the correlation was not another model. It was a test suite, an external check with a completely different failure mode, which is why it carried information the panel did not already have. That is what independence looks like in practice. Not a ninth opinion, but a different kind of evidence: run the code, check that the citation resolves, verify the number against the source.

Accuracy tells us whether a judge fails. Correlation tells us whether the backup fails with it. The pivotal cases tell us whether any of it mattered.

How much independence does the decision deserve

Independence is purchased, not switched on, so the sensible question is proportionality. A model suggesting a better sentence, which a writer accepts or rejects on the spot, needs almost none; the human is the control and a bad suggestion costs a keystroke. A resume screen producing a shortlist nobody revisits needs a great deal, because correlated evaluator bias lands on applicants who never learn it happened and have no route to appeal. A safety monitor supervising an autonomous model needs the most and has the least, since the monitor most likely to be deployed is a copy of the thing it watches.

Agentic systems make this urgent at the wrong moment. The natural response to an agent that might act badly is to put another agent above it, and a third above that. The architecture acquires layers. Whether it acquires controls depends on whether those layers fail differently, and the evidence so far says similar models do not.

The separation of powers

Constitutional systems do not assume officeholders will exercise power wisely because someone asked them to. They divide authority on the assumption that they will not. Safety engineering does not assume a component never fails; it builds redundancy and then spends most of its effort on common-mode failure, because redundancy that fails together is decoration.

AI governance is arriving at the same principle from a new direction and is behind on the hardest part. It has begun asking for independent evaluation. It has not worked out how to tell whether 2 machine evaluators are independent in the sense that matters, which is whether they can be wrong in different ways.

Nine judges sound safer than 1. Apple's experiment suggests the number can be deceptive. Before AI gets its own separation of powers, somebody has to establish that the powers are actually separate.


Further Reading

Earlier in this sequence

Primary sources


AI Assistance Statement ▾
Preparation of this blog entry included drafting assistance from ChatGPT using a GPT-5 series reasoning model. The tool was used to help organize ideas, propose structure, refine language, and accelerate revision. It was also used to assist in identifying image sources and verifying that selected images appear to be released for reuse (for example through public domain or Creative Commons licensing). The author selected the topic, determined the argument, reviewed and edited the text, confirmed image licensing, and takes full responsibility for the final published content.

#AIData #Observation