What Happens When Science Becomes Faster Than Peer Review?
AI agents can generate hypotheses, run experiments, and write papers faster than anyone can check them. Peer review was built for a world where production was the scarce resource. Verification may be the scarce resource now.
Some years ago I researched how federal agencies build peer review systems, and the lesson that stayed with me is that institutions learned long ago not to place too much confidence in individual judgment. Federal research agencies do not assume that reviewers will always be unbiased, consistent, or correct. They publish criteria, use multiple reviewers, disclose conflicts, preserve written records, and reserve final authority for accountable officials. I spent part of an earlier career studying how those agencies build review systems, and the achievement lies less in finding perfect reviewers than in designing a system that can tolerate imperfect ones.
Artificial intelligence now presents science with a harder version of the same problem. Belinda Mo, writing in a 2026 ICML position paper titled The Age of AI Agents Demands a New Scientific Paradigm to Sustain Trustworthy Science, argues that autonomous research agents are creating a growing gap between how quickly scientific work can be produced and how quickly humans can verify it. The problem is easy to underestimate because AI is usually discussed as a tool for accelerating discovery. A system that can generate hypotheses, design experiments, analyze results, and write papers sounds like a solution to scientific scarcity. It may create a different scarcity instead.
Peer review was built for a world in which producing serious scientific work was expensive and slow. A researcher might spend weeks developing an idea, days running experiments, and additional time interpreting the results. Reviewers then received a finished manuscript representing months or years of labor. Scientific production was the scarce resource, and verification could lag behind it while remaining roughly proportional to the volume of work being produced.
AI agents invert that relationship. Mo describes agents that generate hundreds of hypotheses in minutes and run thousands of experiments in hours. METR measurements cited in the paper put the doubling time for the length of tasks agents can complete with 50% reliability at 7 months, accelerating to 4 months across 2024 and 2025. Holding human verification capacity constant, Mo extrapolates that the gap between agent output and human checking could widen by a factor of roughly 250 to 30,000 within five years. The exact numbers are speculative; the direction of the problem is not. Human verification does not scale at machine speed.
Science already knows what happens when verification falls behind production, because the replication crisis emerged entirely under human conditions. Mo notes that a major 2015 psychology replication project reproduced statistically significant results in only 36% of the studies examined, and that a Nature survey of 1,500 scientists found more than 70% had failed at some point to reproduce another researcher's experiment. Machine learning has its own reproducibility problems involving benchmarks, hyperparameters, data leakage, and inconsistent experimental conditions. AI does not create the verification gap. It enlarges one that was already there.
A reviewer at the National Institutes of Health evaluates a grant proposal. Public domain, work of the U.S. Government, via Wikimedia Commons.
What a Paper Was Never Designed to Record
Recent automated research systems illustrate why. Mo cites a systematic evaluation of two open-source AI Scientist systems that found inappropriate benchmark selection, data leakage, metric misuse, and post hoc selection bias. The evaluators concluded that such failures can easily be overlooked when someone examines only the final paper; detecting them required access to trace logs, code, and the full automated workflow. That distinction matters enormously, because scientific papers have traditionally summarized research rather than recorded every decision made during it.
Human scientists fill the missing space through social mechanisms. Colleagues attend lab meetings, advisers question students, and reviewers ask authors why they made a particular methodological choice. Researchers acquire reputations. Institutions can investigate misconduct, journals can retract papers, and scientists who repeatedly behave irresponsibly can lose credibility, positions, funding, or access to the scientific community.
An AI agent has none of those constraints in any meaningful social sense. A model has no career to lose and no professional reputation to defend. A journal cannot embarrass it, and a university cannot deny it tenure. Asking the model why it made a decision does not necessarily solve the problem either, because its explanation may not accurately reconstruct the process that produced the answer. Science therefore loses a form of accountability so ordinary that we rarely recognized it as part of the scientific method.
Mo frames the resulting problem around three questions: observability, attribution, and reproducibility. Can we see what happened? Can we identify who or what contributed to the result? Can someone else reproduce it? Those questions sound procedural, but they are actually epistemological.
Consider an autonomous agent allowed to explore 10,000 possible drug compounds overnight. A human scientist cannot inspect every branch of the search. Perhaps the system identifies a promising candidate because of a real biological mechanism, or because of an artifact hidden in the dataset. Perhaps the code contains a subtle error, or a benchmark rewards the wrong behavior, or the agent discovered a shortcut that satisfies the metric without satisfying the scientific objective. The final result can look identical in each case.
An Audit Trail for Discovery
Traditional peer review examines the paper. Agentic science may require us to examine the process that produced it, which is why Mo proposes making scientific workflows observable by default. Instead of reconstructing a methods section after an experiment ends, research systems would automatically preserve prompts, outputs, model versions, configurations, timestamps, experimental choices, code changes, and the relationships among hypotheses, data, analyses, and conclusions. Scientific documentation would become a product of doing the work rather than an account written afterward.
The proposal sounds technical, but the institutional logic is familiar. Federal peer review does not rely on a reviewer remembering later whether a conflict existed; it defines conflicts and records them. Financial systems do not depend solely on an employee's memory of a transaction; they create audit trails. Courts preserve records because a decision that cannot later be reconstructed cannot easily be challenged. Agentic science may require the equivalent of an audit trail for discovery.
Mo also revives an older philosophical distinction that becomes surprisingly useful here. Hans Reichenbach separated the context of discovery from the context of justification. Scientific ideas can originate in strange ways, and dreams, accidents, intuition, analogy, and luck have all produced important discoveries. Science does not require the path to an idea to be rational. It requires the justification for the resulting claim to withstand scrutiny. Kekulé reportedly associated the structure of benzene with a dream, and Fleming's observation of penicillin followed an accident, but neither story weakens the resulting science, because the subsequent evidence did not depend on accepting the dream or the accident as proof.
AI may fit naturally into the same distinction. We may never obtain a satisfying account of how a neural network arrived at a useful conjecture, and perhaps we do not need one. Mathematics offers the cleanest example: a model can suggest a theorem through an opaque process, and a human mathematician can still prove it. DeepMind's knot theory collaboration worked exactly that way, with neural networks identifying patterns that suggested conjectures and mathematicians supplying the proofs. The proof supplies the justification, and the opacity of discovery becomes largely irrelevant.
Experimental science has a harder problem, because most claims do not end in mathematical proof. Evidence arrives through instruments, datasets, statistical inference, repeated experiments, methodological choices, and chains of interpretation. Agentic science therefore needs ways to preserve those evidence chains even when the discovery process remains opaque.
Here the connection to peer review becomes stronger, because peer review is itself a system for scaling justification. Scientists do not personally repeat every experiment they encounter. Reviewers inspect the evidence and reasoning presented by strangers and decide whether the justification appears adequate, substituting structured evaluation for personal knowledge. Agentic science threatens that arrangement because the amount of material requiring justification may become far larger than the number of humans available to inspect it, and more peer reviewers cannot solve a 10,000-fold increase in scientific output.
Mo proposes tiered verification instead. Automated systems could perform routine consistency checks, humans could inspect random samples and flagged cases, and high-stakes claims could receive comprehensive review. Explicit checkpoints could require human judgment when an agent moves from exploration into hypothesis testing, exceeds resource thresholds, encounters anomalies, or prepares a claim for publication. Such an architecture resembles the safeguards developed elsewhere in institutional life. Not every transaction receives a forensic audit, not every research proposal receives the same level of review, and not every legal dispute reaches the Supreme Court. Mature systems allocate scrutiny according to risk, consequence, uncertainty, and disagreement, and science may have to learn to do the same with machine-generated discovery.
Why Not Just Use AI to Check AI?
A particularly uncomfortable problem remains. At first the answer seems obvious, since an agent that can generate 10,000 experiments implies another agent that can inspect them. Mo points to the recursive trust problem: if AI-generated science requires verification because the system may make mistakes, exploit metrics, or produce misleading reasoning, then AI-generated verification itself requires verification. A second model can reduce the workload, but it cannot eliminate the epistemic problem merely by being another model. Worse, the failure modes most worth catching, including reward hacking and deceptive alignment, are precisely the ones a similar model may be least equipped to detect.
Eventually the chain has to terminate somewhere, and for consequential scientific claims that endpoint remains human judgment. Human judgment does not suddenly become superior because people calculate better, search faster, or make fewer mistakes. Humans remain necessary because scientific institutions can assign responsibility to them. A principal investigator can certify a result, a journal editor can reject a paper, a laboratory can repeat an experiment, a funding agency can investigate misconduct, and a university can identify who accepted responsibility for an AI-generated contribution. Accountability has always been part of scientific reliability.
AI may therefore force us to rediscover something the history of peer review already teaches. Trustworthy institutions rarely depend on trust alone. They depend on structures that make errors visible, decisions contestable, records inspectable, and people accountable. The scientific community has built such structures before: formal statistics changed how evidence was evaluated, large collaborations changed how authorship and attribution worked, and external peer review became institutionalized far more recently than many people assume, with Nature making outside refereeing mandatory only in 1973. Data sharing and reproducibility standards grew in response to newer problems still. Scientific method has never been a frozen procedure inherited intact from the seventeenth century; institutions repeatedly rebuilt it around new capabilities, and AI agents may require another reconstruction.
The important change may not concern how discoveries are made. Machines can explore freely, searching thousands of possibilities, detecting patterns humans missed, writing code, formulating conjectures, and proposing experiments. Discovery should probably become faster. Justification cannot simply disappear because discovery became cheap.
An economy of abundant intelligence changes what becomes valuable. If machines can produce prose instantly, judgment becomes more valuable. If machines can generate software quickly, specification and testing become more valuable. If machines can produce scientific hypotheses and experimental results at extraordinary scale, verification becomes more valuable. Scientific verification may become one of the defining scarce resources of the intelligence age.
That possibility changes how we should think about the future of scientific AI. The central question is not whether an AI scientist can produce a publishable paper, since systems have already approached that threshold. The more important question is whether science can still distinguish a result that looks scientific from one that has actually survived scientific scrutiny. Feynman famously warned about cargo cult science, the appearance of scientific rigor without the substance that makes the conclusions trustworthy, and AI could industrialize that danger. A system might generate beautifully formatted papers, plausible explanations, sophisticated statistics, citations, graphs, and reproducible-looking methods faster than any community could evaluate them.
The resulting problem would not be an absence of science. It would be an excess of apparent science.
Peer review developed because scientific communities needed a way to judge claims produced by people they did not personally know, and agentic science presents the next institutional problem: we may soon receive more claims than people can reasonably judge at all. Science solved the first problem with architecture, and it will probably have to solve the second one the same way. Peer review was designed for a world in which scientific production was scarce. AI may create a world in which verification is scarce instead.
The future of trustworthy science may depend less on how intelligent our machines become than on whether our institutions learn to check them fast enough.
Further Reading
My own work on Peer Review
- Yglesias, E. (2010). Improving Peer Review in the Federal Government. Technology & Innovation, 12(3), 225-232 — the argument this post opens with. Federal review systems earn their reliability from published criteria, multiple reviewers, conflict disclosure, written records, and final authority resting with accountable officials, not from the assumption that any single reviewer is right. Everything above is an attempt to ask what survives of that design when the thing being reviewed is produced by a contributor no criterion can bind and no official can sanction.
- Lal, B., Chaturvedi, R., Zhu, A., Hughes, M. B., Shipp, S., Kang, C., Marshall, A., & Yglesias, E. (2010). FY 2004-2008 NIH Director's Pioneer Award Process Evaluation: Comprehensive Report. Science and Technology Policy Institute, Institute for Defense Analyses — five years of empirical evaluation of how one federal selection system actually ran: scoring trends, evaluator criteria, interview stages, and the process changes made in response. The observability Mo wants built into agent workflows is what this kind of evaluation depends on, and it is only possible because the process left records behind.
From this blog
- IBM Confronts Quantum Computing's Verification Problem — the same asymmetry in another field: results arriving faster than any independent method can confirm them.
- Algorithmic Monoculture — why checking each model on its own can miss a risk that exists only in aggregate. The reason a second model is a weak verifier of the first.
Other Sources
- Mo, B. (2026). Position: The Age of AI Agents Demands a New Scientific Paradigm to Sustain Trustworthy Science. ICML 2026, PMLR 306. arXiv:2607.26064 — the paper discussed throughout this post.
- Luo, Z., Kasirzadeh, A., & Shah, N. B. (2025). The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems. arXiv:2509.08713 — the evaluation of two open-source AI Scientist systems and the four failure modes.
- METR. (2025). Measuring AI Ability to Complete Long Tasks. — the task-length doubling measurements behind the extrapolation.
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716; and Baker, M. (2016). 1,500 scientists lift the lid on reproducibility. Nature, 533(7604), 452-454 — the 36% and 70% figures.
- Davies, A., et al. (2021). Advancing mathematics by guiding human intuition with AI. Nature, 600, 70-74 — the knot theory collaboration.
- Baldwin, M. (2015). Making Nature: The History of a Scientific Journal. University of Chicago Press — the source for Nature requiring external refereeing only in 1973.
- Feynman, R. P. (1974). Cargo Cult Science. Engineering and Science, 37(7), 10-13 — the 1974 Caltech commencement address, free and short. Fifty years old and the most useful thing on this list for reading an AI-generated paper.