elmerdata.ai blog

My blog

AI ♥ AI

A job applicant uses AI to write his resume. The employer uses AI to screen it. Recent research finds self preference rates approaching 90% among some models, and the advantage is large enough to change who gets shortlisted. The loop does not need a conspiracy to close.


Narcissus lying at the edge of a pool gazing at his reflection while the nymph Echo watches from across the water John William Waterhouse, Echo and Narcissus, 1903. Oil on canvas, 109.2 by 189.2 cm. Walker Art Gallery, Liverpool, purchased the year it was painted. Public domain. Echo is the figure in the shadows on the left, and the curse laid on her is that she can only repeat what she has already heard.

Callie Holtermann's New York Times Magazine piece last month opens with a 30-year-old finance candidate named Christian Vinson, who spent about 10 hours teaching 2 chatbots his background and then had them write cover letters for hundreds of listings. AI wrote the applications and AI read them. He got the job at a New York investment bank and called it the new normal. The article quotes Soheil Feizi at the University of Maryland on what worries him about systems built on the same models talking to each other: "They are just sharing the same blind spots."

Sharing a blind spot is the charitable version. The research that has accumulated over the past 2 years says something less comfortable, which is that the machine on the reading side of that exchange does not merely fail to notice machine writing. It likes it.

The measured part

Jiannan Xu, Gujie Li and Jane Yi Jiang at Maryland's Smith School ran the experiment the Vinson anecdote implies. They took 2,245 human-written resumes from LiveCareer, had 9 commercial and open-source models produce rewritten versions of each, and then put the models in the evaluator's chair. Controlling for content quality, they found that every major model preferred its own rewrite over an equally good human original, with self preference rates ranging roughly from 67% to 82% in the researchers' June 2026 summary. They then simulated hiring pipelines across 24 occupations and found that a candidate using the same model as the employer was 23% to 60% more likely to reach a shortlist than an equally qualified candidate who wrote the resume himself, with the widest gaps in sales and accounting. The paper went up as a preprint in August 2025, has been revised through June 2026, and is now in the AAAI/ACM conference proceedings on AI, ethics and society.

The effect is not confined to hiring. Walter Laurito and colleagues published a study in PNAS in July 2025 comparing human and LLM descriptions of the same underlying items, across 109 consumer products, 100 scientific papers and 250 film summaries. When GPT-4 wrote the pitch, human evaluators picked the AI-described product 36% of the time. LLM evaluators picked it 89% of the time. The authors are willing to call that what it looks like, warning of an "LLM writing-assistance tax" on people who cannot afford the tools.

Notice what kind of bias this is. Algorithmic fairness has spent a decade asking whether a system treats people differently by race, sex, age or disability, and those questions are about who the applicant is. What the resume experiment measures is which vendor the applicant used. It is a protected class of one machine, and no existing framework has a name for it.

Why a machine would prefer its own sentences

Start by taking the title away. Affection is the frame everyone reaches for, mine included, and it is wrong: these models are not recognizing kin. The honest mechanism is duller and more interesting. Arjun Panickssery and colleagues showed at NeurIPS in 2024 that a model's ability to recognize its own text correlates linearly with the strength of its self-preference, and that fine-tuning a model to recognize itself more or less moves the preference along with it. Familiarity, not loyalty. A model has an implicit sense of what a sentence should look like, and its own output sits closest to that expectation.

Two recent papers make the picture sharper in opposite directions. Wei-Lin Chen and coauthors argue that much of what looks like self-preference in strong evaluators is legitimate, because the strong model often did produce the better answer, which is a real caution against reading every preference as bias. Their exception is the part that matters: when the evaluator's own output is actually wrong, its self-preference becomes pronounced, and the stronger the model, the harder it finds its own errors. José Pombal, Ricardo Rei and André Martins then tested this against rubrics a computer can verify, using IFEval and LiveCodeBench, where satisfying a requirement is not a matter of taste. Judges can be up to 50% more likely to wrongly mark a requirement as satisfied when the failing output is their own.

That is the finding to keep. The bias is smallest where a model is right and largest where it is wrong, which is precisely backwards from what an evaluator is for.

Data inbreeding

Now run the loop forward. Graphite sampled 43,000 English-language articles from CommonCrawl and found that AI-generated articles passed human-written ones in November 2024 and hovered around half of new output through the study's May 2025 cutoff. Model output is becoming a large share of the text pool that models are trained on, ranked by, and evaluated against.

Ilia Shumailov and coauthors gave the training-side version of this a name in Nature in July 2024. Train each generation on the output of the last and the tails of the distribution disappear first, then the variance, until the model converges on a bland average of itself. They called it model collapse, and the coverage treated it as a prophecy.

It probably is not one, and the correction deserves as much attention as the original. Matthias Gerstgrasser, Rylan Schaeffer and colleagues showed that the collapse result depends on synthetic data replacing real data with each generation. If the data accumulate instead, which is what actually happens on a web that keeps its old pages, the error has a finite ceiling and the collapse does not arrive. So the apocalyptic reading of data inbreeding is weaker than it sounds.

Which is why the evaluation-side finding is the one I would worry about. Self-preference is not a passive contamination of the pool. It is a selection pressure on it. If machine evaluators sit at the chokepoints where resumes become interviews, papers become citations, proposals become contracts and pages become traffic, then machine-flavored text does not simply pile up alongside human text. It wins slightly more often, and everyone downstream learns to produce more of it. Resume writers already optimize for applicant tracking systems and publishers already optimize for search engines, so nobody should expect this pressure to go unanswered.

That last step is an extrapolation and deserves to be labeled one. What the studies measure is a single round of selection, under controlled conditions, at one moment. Nobody has watched the loop run for 10 years and reported back, and the compounding is inference rather than finding. Still, the end state I would bet on is not that people are replaced by machines. It is that people write in a machine-optimized dialect because that is what passes.

None of this requires anyone to decide that AI writing is better. Every party in the chain is behaving reasonably. The applicant uses the best tool available. The employer screens hundreds of applications with the only method that fits the budget. The vendor sells a good product. The loop closes on its own, which is the same structure I traced in algorithmic monoculture and again in when the firm becomes the monoculture. Nobody chooses it and everybody builds it.

The old rule

The good news in the Maryland paper is easy to miss. System prompting the evaluator to disregard the source, and majority voting across models with weaker self-recognition, each cut the bias by more than half. The problem is tractable, and it is a design problem with a known direction of travel, not a law of nature.

But the fix points at something institutions already know. A model checking another model is routinely treated as independent review, and when the 2 come from the same family it is not review at all, it is the author grading his own paper in a different font. Peer review, judicial recusal, external audit, separation of duties and the entire apparatus of institutional checking exist because we worked out a long time ago that evaluation requires distance from authorship. I made a version of this argument about AI recusal and about science outrunning peer review. It applies here with less ambiguity, because here we have the effect size.

So the governance question is not whether AI evaluation is accurate. It is whether it is independent, and independence is now a property of the procurement decision. An organization that buys its writing assistant and its screening tool from the same vendor has not deployed 2 systems. It has deployed 1 system twice and called the second instance a check.

Narcissus was not punished with vanity. He was punished with an inability to recognize that what he was looking at was himself, and he stayed at the water until he wasted away. Our machines have the same difficulty and none of the pathos. They cannot tell their own reflection from the world, and we have started asking them who deserves the job.


Further Reading

From this blog

Primary sources


AI Assistance Statement ▾
Preparation of this blog entry included drafting assistance from ChatGPT using a GPT-5 series reasoning model. The tool was used to help organize ideas, propose structure, refine language, and accelerate revision. It was also used to assist in identifying image sources and verifying that selected images appear to be released for reuse (for example through public domain or Creative Commons licensing). The author selected the topic, determined the argument, reviewed and edited the text, confirmed image licensing, and takes full responsibility for the final published content.

#AIData #Observations