← All Labs · The Paper Route · The stops: 🪑 Park Bench · ⏳ The Docket · 🔨 Sledgehammer Wing · 📐 Caliper Room · 🪟 Teachers’ Lounge
🌲 Opathorlokan University opathorlokanuniversity.net
Essay · The Front Door Methodology & Doctrine · the front door of The Paper Route · §4.7.13 · College 00 Nobody on either side of the handshake can be assumed anymore.
📰 The Paper Route · College 00 · Methodology & Doctrine

The Times Have Changed

How the scientific paper stopped being a finished thing.

Two small failures showed up in the news this year, and on the surface they point in opposite directions.

In the first, security researchers at RWTH Aachen ran the LaTeX source of nearly three million arXiv papers through a scanner — about 2.7 million submissions, reaching back to 1991. Around 88% of the ones that shipped their source files were carrying something the authors never meant to hand over: arguments between co-authors, unflattering comments about rivals, to-do notes acknowledging the weaknesses they’d papered over rather than fixed. And buried in the same source, the study counted 265 API keys, 171 passwords, and four private keys — plus GPS-tagged images that, in a handful of spot-checked cases, resolved to both a research building and a home address. Not “hundreds of secrets.” A precise, un-round tally: hundreds of API keys, yes — but exactly four private keys. The specific number is less quotable than “hundreds” and more damning for it. The clean PDF on top; the whole human mess dragged along underneath, permanently, in public.

In the second, two short papers Max Planck published in a German journal — one in 1940, one in 1942 — spent about fifteen years as blank white pages. Springer Nature retracted both in December 2011, during a mass digitization of its archives. The notices gave one reason: “copyright violation.” No date, no name, no explanation. They sat that way — pages blank, PDFs still on sale — until two historians noticed Planck’s name on a list of Nobel laureates with retractions and wrote it up in 2026.

Here is the part nobody can close: no one can say who did it. The journal’s own editor guessed it was the publisher’s automated policing software. The historians suspect an automated copyright-detection workflow flagged the 1942 essay because it had also appeared elsewhere. Springer Nature flatly denies all of it — its head of research integrity told Retraction Watch the retraction “was a human error” and that “no software or ‘bot’ would have been involved in the process.” The historians don’t buy the denial. And that is exactly where it ends: the people are gone, the records don’t exist, and nobody kept an account of how a Planck paper got erased. On 6 July 2026, Springer Nature quietly reinstated both papers — the blank went back to being a paper by the same untraceable process that made it blank in the first place.

One story is about too much coming out. The other is about something quietly deleted — and just as quietly restored. Revelation and erasure. If you try to weld them at that level, they fight each other.

They don’t fight one floor down.

the shape underneath

The paper is no longer frozen

For most of the history of science, a published paper was frozen ink. The process that produced it — the dead ends, the fights, the “what does this even mean” scrawled in a margin — was invisible and gone. And once it was printed, its past was safe. Nobody could reach into a 1942 journal and un-write a sentence — or, later, write it back in.

Both of these stories are the same event: the paper stopped being that object. It became a database record. And a database record does two things frozen ink never could. It leaks forward — the source file carries every comment and secret behind the clean copy, indefinitely. And it’s editable backward — the record can be reached into and changed after the fact, and, as the Planck case shows, changed again, with no reliable account of who changed it or why. Same underlying fact, two faces. The paper is no longer fixed in either direction.

Notice what each failure actually is. The share-your-source norm was built for transparency and reproducibility. It worked — and in working, it surfaced the human mess the norm never meant to expose. The retraction, whoever or whatever performed it, applied a modern housekeeping rule to a wartime essay. And here the historians make the one point in this whole affair that is firmly on the record: categories like “duplicate publication” and “self-plagiarism” are modern inventions that, in their words, cannot be applied to the past “without distorting the historical record.” Their case is sharper than the notice itself, which only ever said copyright violation — and one of the two papers had never appeared anywhere else, so “duplicate” never fit it at all. A rule that didn’t exist when Planck wrote reached back some seventy years and applied anyway, with no one in the loop to say it’s 1942, that category doesn’t apply. Whether the last click was a script or a tired human barely matters. What matters is the judgment call that nobody made.

Hold onto that phrase — no one in the loop. It’s the whole essay.

the last time the times changed

The check was bolted on

Before you decide the sky is falling, go read who built it. The Royal Society keeps a page on its own history, and it is more radical than anything in this essay.

Philosophical Transactions — “the world’s first and longest-running scientific journal,” in the Society’s own words — launched in March 1665. It was not launched by a society. It was launched by a man: Henry Oldenburg (c.1619–1677), the Royal Society’s first Secretary, who was its publisher and its editor. He ran it off his own European correspondence network, monthly, at one shilling a copy. And the Society is refreshingly blunt about why: Oldenburg spun his contacts and his skills as a linguist and scientific editor into a new form of print “intended to promote the enterprise of early modern science and perhaps make some money on the side.

That is the origin of the scientific paper. One well-connected man with a mailing list and a side hustle. After he died the journal passed, unofficially, to whoever held his job next — Halley, Sloane, others. The Society didn’t take institutional control until 1752, and one of the things that pushed it to act was “a series of harsh satires on the Society and its works.” Embarrassment, in other words. The journal has “operated at a loss for most of its existence.” It went online in 1997.

Now the fact that should reorganize how you think about all of this. Here is the Royal Society, describing the 1830s:

“By the 1830s Philosophical Transactions was facing increased competition for the best scientific papers from commercial journals with a more rapid publication schedule… The Royal Society responded by introducing more rigorous and systematic expert peer review.”

Read the dates. 1665 to the 1830s. The scientific paper existed for roughly a hundred and seventy years before systematic peer review did. Peer review is not the paper’s nature. It is not what a paper is. It is a fix — bolted on, late, by an institution that was losing ground to faster competitors and had just been publicly mocked. The check was invented because the old way of trusting a paper stopped working.

So when someone tells you the current system is the system, that peer review is the immovable floor beneath science: it isn’t, and the people who built it will tell you so on their own website. The trust mechanism has been renegotiated before. It was renegotiated under exactly the conditions we’re living in now — too much material, moving too fast, and a growing suspicion that the old filter couldn’t tell good from bad anymore. The last time the times changed, the answer was invent the check.

Hold that. We’re going to need it at the end.

where the modern version began

Back to the beginning

To see where this goes, start where the modern version started. In 1991, Paul Ginsparg built arXiv so physicists could share preprints without waiting on the slow machinery of journals. It worked because of a specific, unspoken thing: human-to-human trust. An endorsement web, where getting in meant an established author vouched for you. Light-touch moderation by volunteers who could tell real work from noise. The whole architecture assumed a person on both ends of the handshake.

On 1 July 2026, after twenty-five years, arXiv left Cornell and became an independent nonprofit. Two reasons, and they rhyme with our two stories. The first is money: the platform ran about $6.7 million in fiscal 2025 against a $297,000 deficit Cornell could no longer absorb, with roughly a three-year funding runway secured and not much visible beyond it. The second is the flood — AI-generated submissions full of fabricated citations and unverified claims. By October 2025, arXiv had already stopped accepting certain computer-science review and position papers unless they’d first cleared peer review. By May 2026 it was handing out one-year bans when moderators found incontrovertible evidence of unvetted machine generation.

Ginsparg named the real problem better than anyone. Recent developments in AI, he said, have “great promise but also pose an existential threat to the underlying arXiv methodology, relying as it does on the bonds of human-to-human trust.” That’s the exact layer AI dissolves. You can buy detection tooling. You can’t buy back the assumption that there’s a person on the other end.

the volume problem

The flood, measured

This isn’t a vibe anymore, and it isn’t an anecdote. It was measured, at scale, in Science.

On 18 December 2025, Kusumegi, Yang, Ginsparg, de Vaan, Stuart and Yin published “Scientific Production in the Era of Large Language Models.” They collected more than two million papers posted between January 2018 and June 2024 on three preprint servers — arXiv, bioRxiv, and SSRN — covering the physical, life, and social sciences. Then they built a detector: compare text from before 2023, when the authors were presumably human, against text an LLM would write, and you can flag which scientists probably started using one. Then count what those scientists published before and after, and check whether journals accepted it.

Stop for a second on the author list. Paul Ginsparg is on it — the same Ginsparg who founded arXiv in 1991, and who told Cornell that AI poses “an existential threat to the underlying arXiv methodology, relying as it does on the bonds of human-to-human trust.” He is not just the man worrying about the flood. He is a co-author of the paper that measured it, using the archive he built as the instrument. That is not a coincidence. That is the person closest to the machinery going and taking the readings himself.

The readings. Productivity went up, hard. On arXiv, scientists who appeared to be using LLMs posted about one-third more papers than comparable scientists who weren’t. On bioRxiv and SSRN, more than 50% more. And the gain landed most on the people the old system taxed hardest — researchers writing science in a second language. Scientists at Asian institutions posted between 43.0% and 89.3% more papers after the detector says they picked up the tools, compared with similar scientists who didn’t. Yin thinks that’s big enough to move where the world’s science gets done.

The good news, before the bad, because there is good news. When scientists go hunting for work to cite, the study found that Bing Chat — the first widely adopted AI-powered search tool — was better at surfacing newer publications and relevant books than traditional search, which keeps handing you the older, more-cited thing everybody already cites. First author Keigo Kusumegi puts it plainly: “People using LLMs are connecting to more diverse knowledge, which might be driving more creative ideas.” That is a real finding, from the same paper, and it belongs on the page next to the damage. This essay does not get to print only the half that suits it.

Now the finding that should stop you cold — and it isn’t the productivity number.

For human-written papers, complex language is a signal of quality. Big words, long sentences, clearly-argued difficulty: across all three preprint servers, human papers that scored high on writing complexity were the most likely to be accepted by a journal. That’s the proxy the whole enterprise runs on. It’s why a well-written paper gets read and a clumsy one doesn’t. It worked.

But papers that scored high on the same complexity test and were probably written by an LLM were less likely to be accepted — “suggesting that despite the convincing language, reviewers deemed many of these papers to have little scientific value.”

The proxy inverted. The signal that used to mean good now, on machine-written work, points the other way.

That is the whole crisis in one sentence, and it is much worse than “there are more papers now.” The instrument didn’t break. The instrument reversed polarity, and it did not tell anyone. Every editor triaging a pile, every reviewer skimming for competence, every hiring committee weighing a CV, every funder scanning an abstract — all of them have spent their careers reading fluency as evidence. Yin calls it a “disconnect between writing quality and scientific quality,” and says the implications are big: editors and reviewers struggle to identify valuable submissions, and universities and funding agencies can no longer evaluate scientists by their productivity. The tell everyone was trained on now fires hardest on exactly the papers you should be most suspicious of.

And it is not one corner of science. Yin: “It is a very widespread pattern, across different fields of science – from physical and computer sciences to biological and social sciences.”

Here is why I trust this paper more than I trust most papers. The authors, unprompted, put a knife in their own result. Cornell’s writeup: “The researchers caution that the new findings are based solely on observations.” What they want next is “a controlled experiment, where some scientists are randomly assigned to use LLMs and others can’t.” They found a correlation across two million papers, and then they said out loud: this is a correlation. Sequence is not cause. Adoption came first and the output changed — but they will not tell you the adoption caused it until someone randomizes it. Nobody made them say that. They said it first. That is the doctrine of this entire Route, showing up in the wild, in Science, in the paper that hands us our worst number. The people who flag their own causal limit before anyone else can are the people you should read.

So: more gets written, faster, by more people — including people the old language barrier was quietly excluding — while the one heuristic the review layer leaned on has turned around and started lying to it. The rest is my inference, not the study’s finding: reviewers are volunteers and there is no mechanism that manufactures more of them when submissions compound, so AI creeps into reviewing next, because of course it does. Writer, reader, reviewer: all three going machine at once, none of them on the schedule the others planned for.

Yin’s closing line is the one to keep. “Already now, the question is not, have you used AI? The question is, how exactly have you used AI and whether it’s helpful or not.” That is the same question this essay is asking about itself, three sections from now.

both sides of the object

What the paper is about to become

Two forces are converging, and they hit the paper from both sides.

The reader side: cross-corpus search at the speed of light. The paper has always been the unit of reading — you find one, you read it, you chase its citations. That’s dissolving. On 30 June 2026 — the same 48 hours arXiv cut loose from Cornell — Anthropic launched Claude Science, a workbench that pulls more than sixty scientific databases and toolkits into one environment and turns an agent loose across all of them: genomics, proteomics, structural biology, cheminformatics, the literature itself. When the whole corpus is queryable at once, the individual PDF stops being the atom of knowledge. The atom dissolves into the search.

The writer side: the machine does the science and writes it up. At that same launch, the product planned and ran a search for a molecule to stabilize the broken enzyme behind phenylketonuria, a rare genetic disease — screening some 2,200 compounds and narrowing to four candidates, live. Anthropic also said it would run its own internal drug-discovery program aimed at neglected and rare diseases — the AI company doing the science, not just selling the tool. OpenAI has GPT-Rosalind, a model purpose-built for biological reasoning. John Jumper, who shared a Nobel for AlphaFold, left DeepMind for Anthropic. Behind all of it is Dario Amodei’s “compressed 21st century” — the claim that AI could fold 50 to 100 years of biological progress into 5 to 10.

And there’s a frontier past even that: proposals like aiXiv, an open-access ecosystem where research proposals and papers are submitted, reviewed, and refined by human and AI scientists alike. Sharpen it a notch and you can see where it points — the paper written by machine, judged by machine, read by machine. Where, in that loop, is the human?

But look at what Claude Science actually leads with, because it’s the tell. The headline feature isn’t the search or the autonomy. It’s the gate. Every figure it generates ships with the exact code and environment that produced it, a plain-language description of how it was made, and the full message history behind it — an artifact that carries its own provenance. And a separate reviewer agent checks citations and calculations, flagging and correcting errors as it goes. That’s not a research convenience. That’s the structural answer to the trust collapse: if you can’t assume a human is checking, build the checking into the object.

So the real question stops being will the machine write the paper — it already does — and becomes can you trust the gate.

the honest part

Can you trust the gate?

Start with the ceiling, and let me be careful about where the number comes from, because this is an essay about doing exactly that. An earlier draft of this section opened on a tidy figure — the best frontier model, hallucinating citations only a few percent of the time. That number is gone. It came from a page with no peer review, no auditable method, and no publisher, and you can read the whole autopsy at the bottom of this page. It was cut rather than caveated, because a number you have to apologize for is not a number you should be leading with.

So here is the peer-reviewed picture, which is both worse and better grounded. GhostCite, a 2026 benchmark, ran thirteen models across forty computer-science research domains and found citation-fabrication rates from 14% all the way to 95% — and that spread is across the models, not across the fields. The single best model in the set still invented citations about one time in seven. The by-field variance inside one model was just as wild: the strongest model fabricated only 2.6% of citations in one domain and over 52% in another; a different model hit a 100% fabrication rate in the artificial-intelligence field. Models invent DOIs, paper titles, and author names with confident, plausible-looking specificity — the exact failure a tired human is worst at catching. There is no model, at any price, that you can hand a citation list and trust unread.

Now the part that decides it. GhostCite also surveyed the people who are supposed to be the backstop. Among authors, 41.5% said they copy-paste citations straight in without checking. Among reviewers, 76.7% said they don’t thoroughly check references — and 80% said they’d never once suspected a fake reference in a submission. Invalid citations in published papers didn’t hold steady either; the rate jumped about 81% in a single year. And — this is the one to read twice — 74.5% of the researchers surveyed said they don’t believe peer review can catch citation errors at all.

Three-quarters of the people inside the system have already concluded that the human check — the thing the whole edifice rests on — can’t do the one job we’re counting on it for.

Put the two halves together. The best engine still fabricates enough that unverified trust is reckless — and the human who’s supposed to be the backstop mostly isn’t there, and increasingly doesn’t believe the backstop works. Neither end of the old handshake can be relied on to hold.

where it lands

The two honest responses

Here’s the through-line the whole piece has been walking toward. For three and a half centuries — since the first scientific journal in 1665 — the paper was a finished human artifact, and the thing holding the system up was human-to-human trust. Everything breaking right now is one of those two assumptions coming undone. The LaTeX leak breaks private. The Planck erasure breaks permanent. The flood breaks human reviewer. Claude Science and its rivals break human author. Four faces of one shift: the trust layer is dissolving from every side at once.

There are only two honest responses, and both have already shown up. Ginsparg’s is social: rebuild the human handshake by force — endorsement, bans, verification, hold the line at the person. Claude Science’s is structural: stop assuming the person and bake the trust into the object — the figure that carries the code that made it, the citation checked before it ships, a paper that can prove itself.

And we know which one history picks, because history already picked it once. The 1830s were the same shape: too much material, faster competitors, an old filter that could no longer sort good from bad, and an institution getting publicly laughed at for it. The Royal Society did not respond by asking gentlemen to be more careful. It built the check. Systematic expert peer review is not an ancient rite — it is 1830s infrastructure, invented under pressure, roughly a hundred and seventy years after the paper it was invented to protect. The mechanism we are now told is sacred was itself the emergency retrofit of its day. That is the precedent, and it is not a metaphor: when the trust mechanism breaks, you don’t restore the trust. You install a check.

And the numbers settle which one has to win. Trust can no longer be chosen — you can’t fix this by picking the most trustworthy model and relaxing, because even the best model fabricates at rates that demand a check, and the humans who are supposed to check mostly don’t, and mostly don’t think checking would help. If trust can’t be chosen, it has to be built in — enforced whether or not anyone is paying attention. A gate, not a good intention.

Which raises the last, sharpest point. A gate is only as good as its independence. Claude Science’s citation checker is real, but a reviewer agent that runs on the same underlying model that wrote the thing is a strong gate, not yet a fully independent one. The robust version is the one the researchers are already converging on. GhostCite says it plainly — no model is uniformly reliable — so cross-checking across models catches errors any one of them misses; and a dumb, mechanical check (does this DOI actually resolve, does this URL actually exist) catches fabrication no matter which clever model produced it. Seven in ten of the researchers surveyed already want that automated check wired into submission itself. Independent verification beats trusted generation. The gate that checks itself is the beginning. The gate that doesn’t care who’s being checked is the point.

That’s the whole story in one frame. The paper used to hide how it was made and stay put once it was printed. The next paper will show how it was made and never stop moving — and it will have to prove itself at the door, because nobody on either side of the old handshake can be assumed anymore. Not the author. Not the reviewer — three-quarters of them have already told us the review can’t catch what it’s supposed to.

The tools that earn the trust will not be the ones that answer fastest. They’ll be the ones that refuse to close the loop for you — that hand you the provenance and make you look, that won’t auto-fill the last mile, that treat the good-enough answer as the thing most likely to be wrong. The paper that survives this is the one that shows its work and dares you to check it.

The paper was the technology that changed how science got done. Now the paper itself is the thing being changed.

The times have changed.

⚖️ a note on this essay itself

Because it’s the whole point of the room this door opens: before this essay went up, it was run cold through a fact-check — the same discipline The Paper Route is about. It did not pass on the first try, and the corrections made it better.

An early draft said “hundreds of passwords and private keys.” The truth is 265 API keys, 171 passwords, and four private keys — the honest count is more damning than the round one. An early draft said an automated bot erased the Planck paper; the publisher denies it and nobody can prove otherwise, so this version says only what’s known — which turned out to be the sharper story. An early draft ended on a statistic — “87% of researchers say they always verify” — that simply isn’t in the source (the source says 87% use AI tools, a different claim), so the ending was rebuilt on a figure that is real: 74.5% doubt peer review can catch citation errors. And a handful of the tidiest numbers in the draft traced back to a marketing blog and a publisher’s content page; every one of them has now been cut — see below.

12 July 2026 — the repair. The two sections cut at the gate for bad provenance have been rebuilt on real sources, and they are stronger than what they replaced. The “flood in numbers” section — whose figures (6 million articles, 5.5 million, 36% of cancer submissions, a 23–89% output gain) all traced back to a journal publisher’s content-marketing page — was cut entirely and rebuilt on Kusumegi et al., “Scientific Production in the Era of Large Language Models,” Science, 18 December 2025 (source 7). And the essay’s historical foundation — which had been carrying a bare, unsourced “1665” and nothing else — is now a load-bearing section of its own, sourced to the Royal Society’s own account of Philosophical Transactions (source 8). It also changed what the essay argues: the Royal Society records that systematic peer review arrived in the 1830s, which means the paper predates its own check by about 170 years. That fact wasn’t in the draft. It should have been. A blog post became a Science paper; a bare date became a documented 360-year precedent. That is what the gate is for.

And the last correction is the one that stings. An early draft of this essay reported that the best frontier model hallucinates citations about 3% of the time. It didn’t come from a lab. It came from a page published by a digital marketing agency — not peer-reviewed, not published anywhere, no method anyone can audit. The 3.1% wasn’t even a citation figure: it was the floor of a code-reference range (3.1–15.4%), lifted out of its own headline and passed around as a citation-accuracy number. And that page wears a badge. An orange dot and the words “ORIGINAL RESEARCH,” awarded to itself, by itself, on its own blog post. A self-issued badge is not a source. Nobody granted that credential. The dot in front of it exists to make it look like somebody did. And the tell was not buried in a footnote — it was the biggest thing on the page.

Which would be a satisfying place to stop, except for one thing: this site was doing it too. Pages here carried a ● real marker — same dot, same move — used as a credential on pages with no citations standing behind it. Twenty-nine of those badges were removed from this site in July 2026. This essay went looking for the disease in someone else’s page and found it in its own house. That is not a confession offered for credit. It is the evidence for the argument: a badge you give yourself proves nothing, and the person least equipped to catch you doing it is you. Which is exactly why the check has to be independent, and mechanical, and indifferent to whose work is being checked — including the check that reads this page.

If you find something here that’s still wrong, that’s the door working. Tell the builder.

the rooms this door opens

The Paper Route

This essay is the front door. Its thesis is the Route’s thesis — the scientific paper stopped being a finished thing — and the rooms behind it are where a claim gets walked, stop by stop, until it’s settled or it isn’t.

sources

What’s behind the numbers

Every load-bearing fact above, with the source it came from. Where a source is weaker than peer-reviewed, that’s said in the text.

  1. The arXiv source-file leak — 88%, 265 API keys, 171 passwords, 4 private keys, GPS-to-home. Pennekamp et al., “Hidden Secrets in the arXiv,” RWTH Aachen, IEEE S&P 2026 — arxiv.org/abs/2604.20927. Coverage: Nature, 9 Jul 2026 — nature.com/articles/d41586-026-02057-8.
  2. The Planck retraction, denial, and reinstatement. “Springer Nature un-retracts Planck papers, citing ‘human error’,” Retraction Watch, 7 Jul 2026 — retractionwatch.com. Also Inside Higher Ed, 29 Jun 2026 — insidehighered.com.
  3. The historians’ anachronistic-category argument. Gingras & Khelfaoui — arxiv.org/abs/2605.17534.
  4. arXiv’s finances and independence — $6.7M, $297K deficit, three-year runway. “arXiv sets out on its own,” Physics Today, 8 May 2026 — physicstoday.aip.org. Cornell Chronicle — news.cornell.edu.
  5. arXiv’s CS peer-review policy (Oct 2025) and AI-generation bans (May 2026). arXiv blog, 31 Oct 2025 — blog.arxiv.org.
  6. Ginsparg’s “existential threat” quote. Cornell Chronicle announcement (link above, item 4).
  7. The flood, measured — 2M+ papers on arXiv/bioRxiv/SSRN (Jan 2018–Jun 2024); ~one-third more papers on arXiv and 50%+ on bioRxiv/SSRN after apparent LLM adoption; 43.0%–89.3% for researchers at Asian institutions; complex language predicts acceptance for human papers but not for likely-LLM papers; the more-diverse-knowledge finding; the authors’ own observational caveat; all Yin and Kusumegi quotes. Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart & Yian Yin, “Scientific Production in the Era of Large Language Models,” Science, 18 Dec 2025, DOI 10.1126/science.adw3000 — science.org/doi/10.1126/science.adw3000. Plain-language writeup and the source of every quotation used here: “AI gives scientists a boost, but at the cost of too many mediocre papers,” Cornell Chronicle, 19 Dec 2025 — news.cornell.edu. Note: Paul Ginsparg, quoted elsewhere in this essay as arXiv’s founder, is a co-author of this study.
  8. The 170-year gap — Philosophical Transactions as “the world’s first and longest-running scientific journal”; March 1665; Henry Oldenburg (c.1619–1677) as publisher and editor; monthly, one shilling; “perhaps make some money on the side”; the 1752 institutional takeover after “a series of harsh satires”; and the load-bearing fact — “By the 1830s Philosophical Transactions was facing increased competition… The Royal Society responded by introducing more rigorous and systematic expert peer review”; operated at a loss for most of its existence; online in 1997. The Royal Society, “History of Philosophical Transactions” — royalsociety.org/journals/publishing-activities/publishing350/history-philosophical-transactions.
  9. Claude Science — launch, 60+ databases, the PKU demo, the provenance features, the reviewer agent. Anthropic, 30 Jun 2026 — anthropic.com/news/claude-science-ai-workbench.
  10. GPT-Rosalind. OpenAI — openai.com/index/introducing-gpt-rosalind.
  11. The “compressed 21st century.” Dario Amodei, Machines of Loving Gracedarioamodei.com.
  12. aiXiv — human-and-AI open-access ecosystem. arxiv.org/abs/2508.15126.
  13. Citation fabrication — 14%–95% across 13 models / 40 domains; by-field variance; the survey figures (41.5%, 76.7%, 80%, 74.5%); +81% invalid-citation rate; “no model is uniformly reliable”; support for automated DOI checks. GhostCite — Zhang et al., arxiv.org/abs/2602.06718.
  14. The withdrawn hallucination figures — including the “~3%” citation number, which is in fact the floor of that page’s code-reference range (3.1–15.4%) — and the self-applied “● ORIGINAL RESEARCH” badge discussed in the closing note. DigitalApplied, “AI Hallucination Rate Benchmarks · 2026,” 23 Apr 2026 — digitalapplied.com/blog/ai-model-hallucination-rate-benchmarks-2026-study. This is a digital marketing agency’s self-run benchmark. It is not peer-reviewed, not published in any venue, and its method is not auditable. It is cited here as the object of an argument, not as evidence for one. Every figure it fed into an earlier draft of this essay has been removed.