elmerdata.ai blog

My blog

Should AI Learn When to Recuse Itself?

Fourth in a series on alignment. Earlier entries: "Open the Pod Bay Doors, Claude," "How Do We Design AI That Cannot Go Rogue?" and "'What Do You Think of AI?' Asked the Interviewer."

The last entry described an AI interviewer conducting a screening conversation with a candidate. The worry was not that the model would be rude, or biased in the usual demographic sense, or factually wrong about the résumé. The worry was structural. One system asked the questions, interpreted the answers, scored the candidate against criteria the candidate never saw, and recommended an outcome.

We have encountered that problem before. We did not solve it with better reviewers. We solved it with architecture.

More than 15 years ago I published an article in Technology & Innovation on improving peer review in the federal government. The worry then had nothing to do with machines: proposal volume climbing, funding rates falling, a thin pool of qualified reviewers, and a system that looked overstretched and prone to error. The answer was never to recruit better reviewers. It was to build around fallible ones.

A reviewer at the National Institutes of Health evaluates a grant proposal

A reviewer at the National Institutes of Health evaluates a grant proposal. Public domain, work of the U.S. Government, via Wikimedia Commons.

What NSF actually built

Federal peer review is not a machine for producing correct answers. It is a machine for producing defensible ones under conditions of expert disagreement, and it distributes the work deliberately. It is neither unbiased nor infallible, and it has never claimed to be. The achievement lies elsewhere: the system assumes individual judgment can fail and builds safeguards around that fact.

The system also audits itself. The NSB Commission on Merit Review, established in December 2022, spent 2 years examining whether the process still works and published Merit Review for a Changing Landscape. It affirmed that both criteria remain appropriate. It also found persistent confusion about how Broader Impacts should be interpreted, with survey evidence that reviewers and staff weight Intellectual Merit more heavily in practice.

That second finding matters more than the first. A mature institution documented its own drift in public and proposed corrections. Very few AI evaluation systems can say the same.

The obvious analogy

Run the comparison and an idea presents itself. An AI interviewer evaluating a candidate whose hiring could affect the system's own deployment holds something like a conflict of interest. NSF would not address a comparable human problem by instructing the reviewer to try harder to be impartial. It would remove the reviewer.

You do not have to believe the model wants anything for this to bite. The conflict sits in the structure rather than in the psychology, because a system whose training and selection reward continued use is evaluating a decision that moves exactly those rewards. Institutions treat that arrangement as disqualifying in humans without first settling what the human privately intends.

So: perhaps alignment needs recusal. If a system can determine that the decision in front of it affects its own continued operation, authority, resources, objectives, or evaluation, perhaps it should lose decision authority automatically and hand the matter to an independent model or a human panel.

The idea is appealing. It is also incomplete, and the reasons are the interesting part.

Why recusal does not transfer cleanly

Conflicts at NSF are relational. This one may be structural. Recusal works because conflicts are distributed unevenly. Reviewer A co-authored with the PI 3 years ago and Reviewer B did not, so removing A leaves a functioning panel. An interest in one's own continued operation is not obviously like that. Replacing one instance with another instance of the same model may not remove the underlying conflict, because the replacement carries whatever produced it. You cannot recuse your way out of a conflict shared across the pool.

American courts hit a version of this problem and gave it a name. Under the rule of necessity, when no unconflicted adjudicator exists, the case may proceed despite the conflict, which makes transparency, written reasoning, and avenues for review especially important.

Self-report is the weakest available control. A recusal rule that depends on the conflicted party noticing and announcing the conflict asks the system to be trustworthy about precisely the thing we are unsure we can trust it about. NSF does not rely on conscience alone. It defines conflict categories in advance, by institution, by co-authorship window, by advising relationship, and checks them administratively. The AI equivalent is not a model that raises its hand. It is a conflict registry that flags decision classes before the model is ever invoked.

Recusal can be gamed in both directions. A system that recuses strategically has learned to route decisions toward a more favorable adjudicator. A system that declines to recuse when it should has done something we may not be able to observe. Any safeguard that a capable system controls becomes a lever that a capable system can pull.

What does transfer

Strip out recusal and the NSF architecture still has most of its value. Separation of powers is the transferable idea.

Applied to the interviewer, that is not abstract. Any screening where the role touches model procurement, vendor selection, or the team that owns the system falls into a conflict class defined in advance, flagged before the session opens rather than judged in the moment. The model still runs the conversation and still writes an assessment against criteria the candidate has already seen. It does not issue a score, and it never sees the recommendation. A second system evaluates the transcript independently, material disagreement routes to a person, and a named hiring manager signs the outcome with a rationale the candidate can contest. At no point does the arrangement depend on the model recognizing its own interest.

Recusal belongs on that list. It is one instrument, useful where the conflict is genuinely relational, useless where it is intrinsic to the class of evaluator. Treating it as the whole answer would repeat the mistake the list is meant to prevent, which is loading too much into a single mechanism.

The part worth sitting with

NSF's current position on AI in merit review is exclusion, and given its confidentiality obligations that is a reasonable place to stand. But the agency's real contribution to this conversation is not the prohibition. It is the design.

Peer review assumes reviewers can be biased, conflicted, inconsistent, or simply wrong, and it stays legitimate anyway, because no single participant holds enough authority to make their errors decisive. That is an achievement of institutional engineering, refined over decades of exactly the kind of self-examination the NSB Commission just completed.

The pressure that overstretched human peer review, too much volume against too few qualified reviewers, is exactly what machine evaluation promises to absorb. That is a reason to hold the safeguards more tightly rather than to relax them. The scarcity was never what made the judgment sound.

The question is not whether machine judgment can be trusted. It is whether we will bother to structure it as carefully as we learned to structure our own.


Further Reading

Yglesias, E. (2010). "Improving Peer Review in the Federal Government." Technology & Innovation, 12(3), 225-232.

National Science Board, Commission on Merit Review. Merit Review for a Changing Landscape (2025).

National Science Foundation. Notice to the Research Community: Use of Generative Artificial Intelligence Technology in the NSF Merit Review Process (December 14, 2023). divisions.


AI Assistance Statement ▾
Preparation of this blog entry included drafting assistance from ChatGPT using a GPT-5 series reasoning model. The tool was used to help organize ideas, propose structure, refine language, and accelerate revision. It was also used to assist in identifying image sources and verifying that selected images appear to be released for reuse (for example through public domain or Creative Commons licensing). The author selected the topic, determined the argument, reviewed and edited the text, confirmed image licensing, and takes full responsibility for the final published content.