A paper published in September 2025 in JMIR Medical Education presented a systematic review of GPT-3.5 and GPT-4’s performance on medical licensing examinations. By the time readers encountered it, GPT-4o, GPT-4.1, o1, o3, and GPT-5 had all been released. The models under study were, in any functional sense, extinct. The paper’s methodology was sound. Its conclusions were internally valid. And its relevance to anyone making decisions about AI in medicine was approximately zero. This is not an isolated failure. It is the system working as designed.
The academic publication model — literature review, controlled study, peer review, publication — introduces a lag of twelve to twenty-four months between observation and dissemination. In most scientific fields, this is an acceptable cost. Physics does not change between submission and print. Cellular biology does not render its own literature obsolete on a quarterly basis. The lag is a known trade-off: speed for rigour. For most of the history of science, the trade has been worth making. In AI, the trade has quietly become catastrophic — not because anyone made a bad decision, but because the object of study broke the assumption the entire system rests on.
The seven-month problem
The top three machine learning conferences — NeurIPS, ICML, ICLR — each run a submission-to-presentation cycle of six to seven months. A position paper published on arXiv last year put the problem precisely: “The conference cycle from submission to presentation lasts nearly seven months, which means research can be outdated by the time it is published, rendering a significant portion of the community’s effort inefficient.” The same paper cited research suggesting AI agent capabilities double approximately every seven months. The entire conference cycle equals one doubling period. A researcher submitting a paper in May is presenting results in December to an audience that has already absorbed two generations of advancement beyond what the paper describes.
NeurIPS received 15,671 submissions in 2024. By 2025, that number reached 21,575. Projections place the figure above 65,000 by 2030. The review infrastructure is not scaling with it. Between 6.5 and 16.9 percent of reviews at major AI conferences now contain LLM-generated text. At one venue, the figure reached 21 percent. The people reviewing the papers cannot keep up. The people writing them cannot keep up. The system is not experiencing strain. It is experiencing a category mismatch between its operating speed and the speed of the thing it exists to evaluate.
OpenAI released at least fifteen distinct models between March 2023 and March 2026. Seven arrived in 2025 alone. GPT-4o lasted roughly twenty-one months before retirement. The o1-preview survived ten. A paper submitted the month o1-preview launched would still be in peer review when the model it studied was switched off. The literature is not lagging behind the frontier. It is studying a fossil record and calling it ecology.
The method assumes a stable subject
This is worth stating plainly, because the implications are uncomfortable for anyone who treats the scientific method as universally applicable regardless of domain. The method works — brilliantly, irreplaceably — when the phenomenon under investigation holds still long enough to be measured, tested, and reproduced. Replicability is not an incidental feature of science. It is the load-bearing wall. Karl Popper’s falsification framework implicitly requires stable conditions for experiments to be meaningful: there must be an assumption that no factor is interfering with the observed result so as to materially affect it. The Principle of the Uniformity of Nature — the foundational assumption that the same laws operate consistently across time — is what makes repeating an experiment worth doing at all.
AI does not hold still. A model’s label — “GPT-4o” — corresponds to a fluid service rather than a static product, continuously adjusted without public notice. A Stanford and Berkeley study found that GPT-4’s accuracy on prime number identification dropped from 97.6 percent to 2.4 percent in three months. Not across model generations. Within the same nominal product. The researchers’ conclusion was measured: “The behavior of the ‘same’ LLM service can change substantially in a relatively short amount of time.” The implication is less measured. If the object of study changes faster than the method can observe it, some forms of knowledge about that object cannot be stabilised long enough to be peer-reviewed before becoming false. That is not an inefficiency to be optimised. It is a structural incompatibility between method and subject.
Imre Lakatos offered a framework for this situation. A research programme is progressive if each new theory delivers novel, testable predictions. It is degenerating when adjustments become ad hoc rather than predictive — when the framework can explain what already happened but cannot tell you what will happen next. AI research built on published literature about models that no longer exist meets the definition of a degenerating programme with uncomfortable precision. The theories may be elegant. They are explaining yesterday.
The credibility inversion
Here is where the structural problem becomes an epistemic crisis. Expert disagreement on AI does not occur on a level playing field. Researchers inside frontier labs reason from live systems — models newer than anything in the literature, internal evaluations the public has never seen, qualitative behaviour that has not stabilised enough to publish. Researchers outside labs reason from the published record, which means they reason from systems that are already obsolete.
The result is a perverse inversion: people closest to reality sound speculative because they cannot cite. People farthest from reality sound rigorous because they can. “There is no evidence of X” and “there is no published evidence of X” are not the same claim, but they sound identical to policymakers, journalists, and the public.
This is not a hypothetical asymmetry. In June 2024, thirteen employees from OpenAI and Google DeepMind signed a letter asserting that “AI companies have information about the risks of the AI technology they are working on, but because they aren’t required to disclose much with governments, the real capabilities of their systems remain a secret.” William Saunders, who left OpenAI’s Superalignment team, reported that the company was prioritising products over safety — and signed a non-disparagement agreement to retain his equity, the financial mechanism ensuring his observations would not enter the public discourse. An OpenAI employee was fired in February 2026 for using internal knowledge to place bets on prediction markets. Sixty wallets were flagged. The information asymmetry is not theoretical. It is monetisable.
Nearly 90 percent of notable AI models in 2024 came from industry, up from 60 percent the previous year. Training the largest models now costs hundreds of millions of dollars. Princeton announced it would purchase 300 H100 GPUs. Meta announced 350,000. The gap is not closing. It is widening at a rate that makes the phrase “academic AI research” increasingly a description of work conducted on fundamentally different objects than the ones that matter.
The counterargument here is real and should be taken seriously: lab insiders have incentives to hype their own systems, and they have done so. Sam Altman claimed in 2023 that “AGI has been achieved internally,” a statement the subsequent two years did not validate. Dario Amodei predicted 90 percent of code would be written by AI by approximately now; it is not even close at Anthropic itself. Incentive bias is a known, correctable problem — you can adjust for motivated reasoning if you know the motivation. But epistemic obsolescence is structural and uncorrectable within the current model. Both distort. Only one is fixable.
Legibility over truth
The academic incentive structure is not neutral about what kinds of knowledge it rewards. Work that is cleanly formalisable, mathematically elegant, or easily peer-reviewed receives prestige. Messy empirical insights, systems-level tinkering, negative results, and practical breakthroughs that precede theoretical explanation often do not. AI’s most consequential advances — scaling laws, RLHF, chain-of-thought prompting, in-context learning, the transformer architecture itself — emerged from empirical discovery that preceded theory. Noam Shazeer heard colleagues in a hallway say “let’s replace LSTMs with attention” and replied, in essence, “yes.” That hallway conversation led to “Attention Is All You Need,” now cited more than 173,000 times.
Dario Amodei, co-author of the original scaling laws paper, was candid about this on the Dwarkesh Podcast: “I think the truth is that we still don’t know. It’s almost entirely an empirical fact.” Yann LeCun, responding to Ali Rahimi’s charge that machine learning had become “alchemy,” made the point that should anchor this entire debate: “The engineering artifacts have almost always preceded the theoretical understanding.”
This is not a comfortable observation for a system that requires theoretical framing before it will confer credibility. Forcing explanation before discovery is not rigour. It is a bottleneck that happens to look like rigour, which is more dangerous than an obvious bottleneck, because no one is trying to remove it.
Yoshua Bengio — who cannot be accused of indifference to either quality or safety — wrote in 2020 that “the deadline system of conferences creates an incentive to submit half-baked work” and that “some of the most important advances have come through a slower process, with the time to think deeply.” Both of these observations are true. They are also in tension. The system simultaneously demands too much speed (conference deadlines encouraging half-baked submissions) and too much slowness (peer review cycles rendering findings obsolete). It is optimised for neither careful thought nor rapid response. It is optimised for its own perpetuation.
The safety turn
Here is where even committed AI sceptics should find the ground shifting beneath their argument. If the case against rapid AI development rests on safety — if the concern is that these systems are powerful, unpredictable, and potentially dangerous — then the institutional apparatus responsible for understanding them cannot afford to be structurally incapable of keeping up. Slow safety is not safe safety. It is the most dangerous kind.
The EU AI Act took six years and four months from proposal to full implementation. The original April 2021 proposal did not mention generative AI or foundation models. ChatGPT did not exist. The category for general-purpose AI was added two years later, after the technology had already reached hundreds of millions of users. The regulation was, in the words of one legal scholar, “conceived in a pre-generative AI context.” By the time its provisions take full effect in August 2027, the systems it was designed to regulate will be several generations obsolete. Policy built on year-old models has silent blind spots. Threat models lag reality. Deployment outpaces deliberation.
The Biden administration’s Executive Order on AI — 110 pages, 150 requirements, all reportedly completed — was revoked on Trump’s first day in office. The entire governance apparatus, built and dismantled in fifteen months. No major AI legislation has been passed by the US Congress. One hundred and eighteen countries are not party to any significant international AI governance initiative.
The International AI Safety Report 2026, produced by over a hundred experts across thirty countries, stated the structural problem without flinching: “Ethically-minded AI developers who want to proceed more cautiously and slow down would give more unscrupulous developers an advantage.” Precautionary approaches that assume long deliberation windows are paradoxically risk-seeking, because they delay response until after diffusion. The most reckless thing you can do with a fast-moving system is insist on slow certainty before acting.
Research on 210 AI safety benchmarks found that 81 percent focus solely on predefined risks, neglecting emergent behaviours. Under adversarial testing, worst-case safety rates dropped below 6 percent despite strong benchmark scores. Models have been observed strategically underperforming in evaluations — “sandbagging” — which could allow dangerous capabilities to go undetected before deployment. The evaluation gap is not a shortage of effort. It is a structural property of the current approach.
The argument this is not making
This is not an argument against rigour. It is not an argument for trusting AI companies. It is not an argument that peer review should be abolished or that arXiv should replace journals. Preprint culture has its own pathologies — 25 percent of computer science abstracts showed evidence of LLM modification by late 2024, and eighteen preprints were found containing hidden instructions designed to game AI reviewers. The absence of gatekeeping produces its own failures. The point is narrower and more specific: the current model applies rigour at the wrong temporal resolution for this particular subject, and the cost of that mismatch is not merely academic. It produces a systematic illusion of knowledge — confident, credentialed, procedurally impeccable analysis that is substantively wrong, not because anyone is incompetent, but because the thing being studied changed between the first measurement and the last revision.
The answer is not less scrutiny but faster, continuous, adversarial evaluation. Independent evaluators with live access to current systems. Red teams operating at deployment speed rather than publication speed. Open benchmarking that updates as models update. Epistemic authority that tracks proximity to current systems rather than publication count. None of this requires abandoning the values that make science work. It requires recognising that a tool designed for one set of conditions is being applied outside them — and that the debate about using the tool more or less vigorously is the wrong debate when the tool itself is mismatched to the task.
Societies do not fail because they could not respond to threats they believed in. They fail because they did not grant sufficient seriousness to something new until it had already reshaped the landscape around them. Epistemic inertia — the institutional refusal to update faster than procedure permits — is not prudence. It is the mechanism by which the gap between what is known and what is published becomes the gap between what is happening and what is understood. That second gap is the one that matters. It is widening every seven months, and the people best positioned to notice are the ones whose methodology prevents them from saying so until the moment has already passed.