elmerdata.ai blog

My blog

Algorithmic Monoculture

Nobody chooses monoculture. Everyone merely chooses the locally sensible option. As AI comes to write the evidence and judge it, hiring, markets, and public discourse may be losing the independent judgment they were built on.


Two things are true about the job market in 2026, and they are really the same thing. Applicants use AI to write résumés. Employers use AI to read them.

The scale is not marginal. A June 2026 survey of 1,500 US hiring managers by Resume Genius found that 58% now use AI to screen applications, up from 35% a year earlier. Treat the number with the skepticism a vendor survey deserves, but note what else it found: heavily AI-generated application materials now rank as the second most common red flag, cited by 49% of the same managers. People who screen with machines penalize applicants for writing with them.

Jessica Grose's recent Times opinion video calls this Spy vs. Spy, and the analogy holds as far as it goes. Employers add screens, applicants optimize against them, employers add more screens.

But the arms race is the surface. Underneath sits a question that framing hides: what happens when the same kind of intelligence writes the evidence and judges it?


Herd of Dülmen wild horses at Merfelder Bruch, Germany
A herd of Dülmen wild horses at the Wildpferdebahn in Merfelder Bruch, Dülmen, North Rhine Westphalia, Germany, photographed by Dietmar Rabich on May 31, 2014. The tightly grouped herd offers a visual metaphor for algorithmic monoculture, where independent actors may begin moving in the same direction. Photograph: Dietmar Rabich, “Merfeld, Wildpferdefang, 2014, 0639,” via Wikimedia Commons. Copyright © Dietmar Rabich. Licensed under Creative Commons Attribution ShareAlike 4.0 International, CC BY SA 4.0. The image is not in the public domain.

The result nobody expects

The concept has a name and a formal treatment. In 2021, Jon Kleinberg and Manish Raghavan published "Algorithmic Monoculture and Social Welfare" in PNAS, defining monoculture as the state in which many decision makers rely on the same algorithm. Their result is the interesting part.

Take 2 firms hiring from the same candidate pool. Each can use its own noisy human evaluator or a shared algorithm that is more accurate than either evaluator. For a range of accuracy levels, adopting the algorithm is a strictly dominant strategy. Yet social welfare, the total value of the hires made, can end up lower than if both firms had used their worse, independent judgment. Kleinberg and Raghavan call it a Braess's paradox for screening decisions: adding a better option makes the system worse.

The mechanism is not bias, and it is not failure. It requires no shock, no shared blind spot, no defective model. It is also not universal, and the authors are careful about this: whether the paradox appears depends on structural properties of the ranking model, and they identify families where monoculture has no effect at all. Where it does appear, it needs nothing more than everyone using the same ranking, so that the firm hiring second is choosing from a pool already sorted by the logic that picked first.

What it looks like in an actual labor market

Until recently that argument had little deployed evidence behind it. At FAccT in June 2026, Rishi Bommasani, Sarah Bana, Kathleen Creel, Dan Jurafsky, and Percy Liang presented "Algorithmic Monocultures in Hiring," using data from the assessment vendor pymetrics: 4,197,168 applications from 3,372,132 applicants to 1,746 positions across 156 employers. Of the applicants who applied to 10 positions, 4% were recommended for rejection from every one.

That sounds unremarkable until you compare it to the right baseline. Some applicants get rejected everywhere by chance. What the authors find is that the rejection rate decays as applicants apply more widely, but more slowly than independence predicts. To push it below 0.1%, an applicant would need to submit 25 applications instead of the 10 that independent decisions would require. And rejection at one employer is not merely correlated with rejection at the next. Sometimes it is mechanical, because 42 of the vendor's models screen applicants at more than one company.

Then the authors run the counterfactual that matters. They sample 939 applicants and score each against all 495 applicable models. Not one is rejected by every model. The applicants who get rejected everywhere are not failing a test of quality. They are encountering the same judgment repeatedly and mistaking it for the market's verdict.

One caveat the authors raise themselves: this is game-based assessment, not résumé screening, and the findings may not generalize to every algorithmic hiring tool. That is a real limit. It is also the first study of its kind, because almost no vendor gives researchers this data.

Bias audits ask a different question

New York City has regulated this space since Local Law 144 took effect in 2023. Covered automated employment decision tools require an independent bias audit within the prior year, publication of a summary, and 10 business days of notice to candidates.

Enforcement has been thin. A State Comptroller audit published in December 2025 found the Department of Consumer and Worker Protection had received 2 complaints in 2 years, and that in a survey of 32 companies where the department identified 1 compliance problem, the auditors identified at least 17.

Set enforcement aside, though, because the deeper issue is conceptual. A bias audit asks whether a tool treats demographic groups differently. It does not ask how many employers use that tool. Those are unrelated questions, and a model could pass its audit at every company that deploys it while producing exactly the concentration the FAccT authors measured. That paper makes the point in its third recommendation: agencies should monitor algorithmic monoculture as a distinct risk. Nobody currently does.

You cannot buy independence

The obvious institutional response is to use a different vendor than everyone else. That may not work.

At ICML 2025, Elliot Kim, Avi Garg, Kenny Peng, and Nikhil Garg evaluated more than 350 language models for correlated errors. On one benchmark, 2 models that both got a question wrong gave the same wrong answer 60% of the time. Shared architectures and shared providers drive some of that. But the correlation persists across models from different providers with different architectures, and it is strongest among the largest and most accurate models. They demonstrate the effect on a résumé screening task specifically.

Diversity of vendor is not diversity of judgment. Two procurement decisions can buy you one opinion.

The evidence converges too

Classical monoculture research studies many decision makers sharing one algorithm. Generative AI adds a second mechanism: the material being evaluated may itself be converging.

At ICLR 2024, Vishakh Padmakumar and He He ran a controlled experiment on argumentative essays. Writers using InstructGPT produced work that was measurably more similar to other writers' work, with lower lexical and content diversity. Base GPT-3 did not produce the effect, and the homogenization traced to text the model contributed rather than to changes in how people wrote.

Applied to hiring, that closes a loop. Any screening criterion becomes weaker once applicants can optimize against it cheaply, which is Goodhart's law doing ordinary work. Generative AI collapses the cost of that optimization to nearly zero. The system stops measuring candidates and starts measuring their access to the same tool.

AI could make every résumé better and the résumé itself worse.


The same shape in markets and in speech

The pattern is not confined to hiring. In April 2025, the Bank of England's Financial Policy Committee warned that AI-driven trading raises "the potential for AI-based participants to take increasingly correlated positions," driven by "the widespread use of a small number of open-source or vendor-provided models or underlying data sets, or a more general convergence on very similar model designs across the market." The BIS Annual Economic Report in June 2026 put it more bluntly: supervisory visibility needs to improve enough to capture model-driven correlated exposures, herding, and possible algorithmic collusion.

I wrote here in August about the SEC's silence on agentic trading. This is the layer underneath that question. Speed amplifies mistakes. Monoculture correlates them. One AI trader making a strange decision is a trading problem. Thousands making the same strange decision at once is a market problem.

The version I find hardest to dismiss involves writing. Stratis Tsirtsis and colleagues at the Hasso Plattner Institute and the Oxford Internet Institute published work in May 2026 finding that language models introduce directional shifts in the positions expressed in text they edit, even when instructed to preserve meaning, then modeled how small per-edit biases compound across a network. Read it carefully: the systematic tests covered 4 open-weight models, and the aggregate effect is simulated rather than observed. The direction is still worth sitting with. We spent 15 years worrying that algorithms would decide which opinions we see. The stranger possibility is that they increasingly help us decide how to say them.


Where the argument is weakest

The literature is not settled, and the 2026 work cuts against alarmism.

Brian Hedden and Manish Raghavan argue that most objections to monoculture are weaker than critics assume. With a fixed number of jobs, the same number of people go unhired either way, and polyculture can favor applicants wealthy enough to apply more widely. The objection they find most compelling is narrower, that greedy shared decision-making under-explores uncertain candidates and locks in what it already believes, and even that one they judge non-decisive, since a monoculture can build exploration into itself.

Robert Kleinberg, Erald Sinanaj, and Éva Tardos went further, quantifying the welfare loss the 2021 paper left open and proving a tight bound of 2 on the price of anarchy. Monoculture is costly, but the cost is capped rather than catastrophic.

And the most-cited empirical result deserves a correction. Bommasani and coauthors' 2022 NeurIPS paper is routinely summarized as showing that shared foundation models homogenize outcomes. It does not. Sharing training data reliably increases homogenization; sharing a foundation model produced mixed results that depended on how the model was adapted. The mechanism is data and deployment, not the base model alone.

None of that dissolves the problem. It reframes it. Monoculture is not proof that common algorithms make every system worse. It is a warning that evaluating each model on its own may miss a risk that exists only in aggregate.

Accuracy is not resilience

Organizations evaluating AI ask about accuracy, bias, hallucination, latency, and cost. Every one of those questions is about a model in isolation. None asks how correlated this model's judgment is with every other model in the same system. That is the question monoculture adds, and it has an uncomfortable property: no individual buyer can answer it, because the answer depends on what everyone else bought.

The agricultural metaphor earns its keep here. Monoculture is efficient and fragile in one specific way: a single failure mode reaches everything at once. Human judgment is inconsistent, and we usually treat that as a defect. Sometimes it is the redundancy. Different recruiters notice different things. Different investors read the same disclosure differently. That variation is not free and not always good, but it keeps one mistake from becoming everyone's mistake.

The central question is no longer whether AI makes good decisions. It is what happens when everyone asks similar AIs to make them.

The safest AI future may not be one in which every institution has the best model. It may be one in which no single way of thinking becomes everybody's model.


Further Reading


AI Assistance Statement ▾
Preparation of this blog entry included drafting assistance from ChatGPT using a GPT-5 series reasoning model. The tool was used to help organize ideas, propose structure, refine language, and accelerate revision. It was also used to assist in identifying image sources and verifying that selected images appear to be released for reuse (for example through public domain or Creative Commons licensing). The author selected the topic, determined the argument, reviewed and edited the text, confirmed image licensing, and takes full responsibility for the final published content.

#AIData #Observations