Six Times the Rate: How AI Is Corrupting the Scientific Record
The Lancet named the velocity. The publication record has the tooling. It does not have the discipline.
Topaz named it
In May 2026, Maxim Topaz at Columbia University and colleagues did something the field had been refusing to do out loud. They audited 2.5 million biomedical papers and named the rate at which fabricated citations were entering the published scientific record.1
In 2023, one in 2,828 biomedical papers carried at least one fabricated reference. In 2025, one in 458. In the first seven weeks of 2026, one in 277.
A six-fold rise in two years. Then, in seven weeks of new data, the rate accelerated again.
The mechanism is not a mystery. Large language models, used to draft scientific papers, hallucinate citations. They generate plausible-sounding author names attached to plausible-sounding titles in plausible-sounding journals, none of which exist. The hallucinations make it through drafting, through co-author review, through journal submission, and through peer review into the published record.
Once a fabricated citation is in a published paper, it acquires the credibility of the publication venue. Subsequent authors cite the original paper, which cites the fabrication. The error compounds.
Topaz did not just measure the rise. He named it. That act of naming, from inside the publication system itself, is what the record has been waiting for. The numbers were always available to anyone who looked. The discipline to look had not been organized.
The publication record has no CASP
The companion essay to this one argued that cancer has more data than protein structure prediction ever had at the moment of AlphaFold’s breakthrough, and that what cancer lacks is the CASP-equivalent infrastructure of disciplined, public, falsifiable tournaments that turned protein data into protein-structure discovery.2
The publication record has the same shape of gap, one layer up the stack.
The infrastructure exists. Crossref catalogs every DOI ever issued. PubMed indexes every biomedical paper. Journal-side databases verify bibliographic data on demand. The tooling to automatically check whether a citation refers to a real paper is technically available and has been for years.
What does not exist is the institution that disciplines the use of that tooling. There is no recurring public tournament in which a journal’s citation hygiene is audited, scored, and graded against a published methodology that survives independent replication. There is no CASP for citations.
The cancer-AlphaFold has been waiting for the discipline that converts the Cancer Research Data Commons into a falsifiable public tournament. The scientific record has been waiting for the discipline that converts Crossref into a publicly-graded defense of what gets into print.
Both are infrastructure-present, discipline-absent. The corruption rate is what happens when infrastructure exists without the discipline to wield it.
The NeurIPS audit
In January 2026, GPTZero researchers published an audit of NeurIPS 2025, the premier machine-learning conference.3 They scanned 4,841 accepted papers. They found 100 hallucinated citations across roughly 53 of those papers, representing about one percent of accepted submissions.
GPTZero coined a term for what they were finding: vibe citations. Citations that match the format of real citations in every dimension except being real. Fabricated author names attached to fabricated titles in real-looking journals, indexed nowhere because they were never published.
Each NeurIPS submission undergoes review by three or more expert reviewers. The reviewers did not catch the fabrications.
This is the most important fact about hallucinated citations. They are not obviously wrong.
A reader scanning a bibliography sees reasonable-looking entries and trusts the format. Format-checking is what peer review is structurally optimized to do. Provenance-checking is not.
Mohammad Hosseini at Northwestern University named the upstream condition when the Topaz study was released. The first part of his diagnosis was infrastructural: there are “people who don’t even want to spend half an hour to check the references of a paper.”4 The second part was deeper. Citation practices themselves have changed with generative AI. The reflective process of reading a paper, taking notes, deciding whether a source is relevant, has been replaced by prompting a language model and accepting whatever it returns. “Engagement with the literature is becoming increasingly more superficial.”
The discipline that defended against fabrication in the pre-LLM era assumed that fabricating a citation required deliberate intent. Authors who would not invent a paper out of malice were trusted not to invent one by accident. That assumption has expired. The system has not been updated.
ICLR 2026 was worse
In March 2026, GPTZero audited ICLR 2026, another premier machine-learning venue. They sampled 300 papers. Of those, they flagged 90 as containing at least one citation that did not appear to exist online. After human verification, 50 of those flags resolved into confirmed hallucinations.5
Flagged at 30 percent before human review. Confirmed at 16.7 percent after.
Roughly one in six sampled submissions to a top conference contained verified hallucinated citations, with three or more expert reviewers per paper.
The trend line is wrong. NeurIPS 2025: about one percent of accepted papers. ICLR 2026: 16.7 percent of sampled papers. The hallucination rate is not stable. It is accelerating, because LLM-assisted drafting is accelerating, and the reviewer-side defense has not adapted.
The ICLR program committee’s own retrospective on the 2026 review process acknowledged the pattern. Hallucinated references contributed to a “relatively high desk rejection rate” compared to past years, though specific counts were not itemized.6 The discipline that exists today is reactive. A submission gets flagged because a reviewer happens to notice; the flag triggers desk rejection; the rejected paper goes elsewhere. There is no upstream institution that prevents the corruption from being submitted in the first place.
Why peer review does not catch this
Reviewers do not check citations. They assume good faith.
This is not a flaw in any individual reviewer. It is a structural property of how peer review works. A reviewer’s job is to evaluate the scientific contribution: the experimental design, the analysis, the claims, the interpretation. Citations are supporting infrastructure. Reviewers spot-check the most relevant citations; they do not verify the bibliography in full.
Full bibliography verification, confirming each citation exists, is correctly attributed, and supports the claim made in the paper, is unpaid additional labor. No reviewer is compensated for it. No reviewer would accept the assignment if they were.
The infrastructure for citation verification exists at the publisher level. Crossref API calls confirm DOI validity in milliseconds. ScienceDirect has begun piloting publisher-led verification tooling. The decision to deploy it at the scale the LLM-corruption rate requires is editorial, not technical.
The discipline gap has the same shape as the discipline gap in AI-for-science discovery. Infrastructure exists. The discipline that converts infrastructure into systematic defense has not been built at scale.
Why this is a cost nobody priced
The scientific record is a public good. It is the substrate on which subsequent research builds. When the record is corrupted, the corruption propagates.
A pharmaceutical company designs a clinical trial based on prior published evidence. If a fabricated citation appears in a key prior paper, the trial design absorbs the fabrication. The downstream cost arrives years later in a Phase II readout that does not match the expected mechanism.
A regulatory body evaluates a new medical AI tool against the published evidence base. If fabricated citations contaminate that evidence base, the regulator’s assessment absorbs the fabrication. The downstream cost arrives in a market authorization that should not have been granted.
A grant committee evaluates a proposal’s literature review. If the proposal cites fabricated work, the committee’s assessment of the proposal’s foundation absorbs the fabrication. The downstream cost arrives in funded research built on a foundation that does not exist.
Each fabrication, once published, becomes background contamination in the field. Removing it requires retraction, which is rare. Tracking downstream propagation requires citation-graph analysis that no journal performs systematically.
This is happening before AI has produced the cancer cure the public conversation is about. The Hassabis-Davos-acknowledged hypothesis-generation problem is unsolved. Meanwhile, the same AI is corrupting the publication record that any future solution would need to be evaluated against.
The promise was that AI would accelerate science. The current measurable effect is that AI is making the scientific record less reliable while not yet producing science of its own.
The Nature Medicine editorial
In April 2026, Nature Medicine published an editorial that named the broader pattern.7
“Evidence that AI tools create value for patients, providers or health systems remains scarce, yet claims about clinical impact are increasingly more common, even though there is no clear agreement on what level of evidence should be required before such claims are considered credible.”
This is what evidence-corruption looks like at the layer above the citation problem. Citation fabrication is the input-side corruption. Inconsistent benchmarking is the output-side corruption. Both make it harder for any reader of the literature to know what is true.
The Nature Medicine venue matters. This is not a critic blog or a position paper from a skeptic group. This is the journal that publishes the medical-AI papers being critiqued, naming the evidence problem in its own primary venue.
When a journal editorializes against its own field’s evidence quality, the field has a structural problem that its publication norms cannot fix from inside.
The discipline that needs to exist
The cancer-AlphaFold needs a CASP. The publication record needs the analog: a discipline that operates at the citation layer, verifies every source before publication, and is funded as a structural role rather than expected as unpaid reviewer labor.
The components of this discipline are not exotic. They exist as scattered features in various publication systems and have never been integrated.
Mandatory citation-verification at manuscript submission. Every cited paper is automatically checked against the DOI registry, verified to exist, and confirmed to match the bibliographic data the author provided. The Crossref infrastructure makes this technically straightforward. The decision to require it is editorial, not technical.
Reviewer-incentive restructuring. Reviewers who catch fabrications get visible credit on their academic record. Currently, reviewing is unpaid and largely unrecognized work; rewarding integrity catches would convert citation verification into a positive-sum activity reviewers have a reason to perform.
Author affirmation of citation review. A required disclosure at submission time that every citation has been manually verified by a named author, with that author taking explicit responsibility for any subsequently identified fabrications. The accountability has to land somewhere.
AI-generated content disclosure. Mandatory and granular: which sections of the manuscript were AI-drafted, which were AI-edited, which were AI-citation-assisted. This is the equivalent of conflict-of-interest disclosure for the LLM era.
A public citation-hygiene tournament. The CASP-analog. A recurring audit, run by an independent body, that scores a journal’s citation hygiene against a published methodology and publishes the results. Journals that participate carry the score on their masthead. Journals that decline carry the absence.
None of these proposals are technically hard. All of them are politically expensive at the publication-system level. They cost the journals revenue, the reviewers time, the authors convenience, and the publishing platforms a marginal user.
That is the structural problem. The discipline that defends the scientific record is unfunded. The corruption it would defend against is free.
For editors and reviewers, this quarter
If you sit on a journal’s editorial board, on a peer-review panel, or on a publishing platform’s product team, three concrete moves are available without waiting for the field to coordinate:
One. Add a citation-verification gate at submission. The technology to automate citation-existence checking exists. Crossref API calls confirm DOI validity in milliseconds. The decision to require it is editorial. Implement it as a submission-system precheck; reject manuscripts whose citations fail verification. Authors will adapt.
Two. Make citation-fabrication a retractable offense. Currently, fabricated citations, when caught, are treated as careless errors. Designate them as research-integrity violations equivalent to data fabrication. The disincentive structure then matches the corruption’s downstream cost. Publish the policy. Apply it to the next year’s submissions.
Three. Publish your AI-disclosure policy. Make explicit which AI-assistance practices require disclosure, which require justification, and which are grounds for desk rejection. The current ambiguity is what permits the corruption rate to accelerate.
These are journal-side moves. They do not require regulatory action, do not require funding agencies to coordinate, and do not require the AI vendors to change anything. They are within the existing authority of the existing actors. They have not been taken at scale.
The close
The scientific record is a public good. AI did not ask permission to corrupt it. The discipline to defend it has not arrived.
The discipline gap in AI-for-science discovery has a sibling here at the publication-record layer. Both are structural. Both have known fixes. Neither has been built at scale because the discipline is unfunded and the corruption is free.
The four Nobels that the AI-for-science conversation invokes (CRISPR, mRNA vaccines, checkpoint inhibitors, AlphaFold) show what scientific discipline looks like in retrospect, after the discovery has been won. The hallucinated citations show what its absence looks like in real time, before the discovery has even been attempted. The same field is doing both.
The pattern generalizes. Wherever a discipline accumulates evidence that subsequent work builds on, the question of who defends the accumulation is the structural question. Clinical guidelines built on retracted papers, medical decisions made against a contaminated literature, regulatory submissions citing work that does not exist: the publication record is one substrate where the discipline gap is visible. There will be others.
The companion essay to this one ended with the line, buy the discipline that runs models, do not buy the models that promised to be the discipline. The version for the publication record is the same shape.
Fund the citation-verification tournament that no journal currently runs. Make the discipline visible. Build the institution that defends what has already been won. Do not wait for the corruption to mandate the discipline.
Topaz did the naming. The next decade of scientific discovery depends on whether anyone funds the response.
---
This is the companion essay to [Superintelligence Won’t Cure Cancer. Discipline Might.](https://appliedsymbioticintelligence.substack.com/p/superintelligence-wont-cure-cancer) The discipline gap in AI-for-science discovery has a sibling at the publication-record layer. Both essays argue that the missing layer is methodology, not intelligence. The cancer-AlphaFold needs a CASP. The publication record needs the analog.
Ryan Gruzen writes Applied Symbiotic Intelligence, on what makes scientific discovery, clinical care, and collaboration actually work, and what makes the patterns last.
If this argument is useful, subscribe and send it to one editor, reviewer, or research-integrity colleague who would recognize the problem.
Topaz, M. et al. *Fabricated citations: an audit across 2.5 million biomedical papers.* The Lancet, May 2026. PII: S0140-6736(26)00603-3. https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(26)00603-3/fulltext. Primary aggregator coverage: Megan Molteni, STAT News, May 7, 2026 (https://www.statnews.com/2026/05/07/lancet-study-finds-steep-rise-fraudulent-citations-academic-papers/). Topaz’s dashboard documenting the study: https://www.maxtopaz.com/citadel.
Gruzen, R. *Superintelligence Won’t Cure Cancer. Discipline Might.* Applied Symbiotic Intelligence, May 13, 2026. https://appliedsymbioticintelligence.substack.com/p/superintelligence-wont-cure-cancer.
GPTZero. *NeurIPS 2025 Audit.* https://gptzero.me/news/neurips/. Academic write-up: Ansari, S. *Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025.* arXiv:2602.05930, February 5, 2026. Media coverage: Fortune, January 21, 2026, https://fortune.com/2026/01/21/neurips-ai-conferences-research-papers-hallucinations/.
Hosseini, M. Northwestern University. Quotes surfaced via Megan Molteni, STAT News, May 7, 2026 (https://www.statnews.com/2026/05/07/lancet-study-finds-steep-rise-fraudulent-citations-academic-papers/). The argument is developed at length in Resnik, D.B. and Hosseini, M. *Hallucinated citations produced by generative artificial intelligence may constitute research misconduct when citations function as data in scholarly papers.* Accountability in Research, March 15, 2026. DOI: 10.1080/08989621.2026.2645390 (verified at Crossref registry; published open access under CC BY 4.0).
GPTZero. *ICLR 2026 Audit.* https://gptzero.me/news/iclr-2026/. Methodology: 300 sampled submissions; 90 flagged pre-human-verification; 50 confirmed as hallucinations after human review (16.7 percent confirmed rate).
ICLR 2026 Program Committee. *A Retrospective on the ICLR 2026 Review Process.* March 31, 2026. https://blog.iclr.cc/2026/03/31/a-retrospective-on-the-iclr-2026-review-process/. 19,525 valid submissions; 779 desk rejections; hallucinated-reference detection cited as contributing factor to elevated desk-rejection rate.
Nature Medicine editorial board. *Show us the evidence for the value of medical AI.* Nature Medicine **32**, 1163 (April 2026). DOI: 10.1038/s41591-026-04389-4 (verified at Crossref registry). Companion piece: Omar, M. et al. *How to meaningfully evaluate AI in clinical medicine.* Nature Medicine (Letter), April 23, 2026. DOI: 10.1038/s41591-026-04350-5 (verified at Crossref registry and PubMed, PMID 42026262).

