← All Labs · The Paper Route ·The Stops · Related: 📰 The Times Have Changed · 🪑 The Park Bench · ⏳ The Docket · 🔨 The Sledgehammer Wing · 📐 The Caliper Room
🌲 Opathorlokan University opathorlokanuniversity.net
Room-hub · Can You Run It Methodology & Doctrine · can you run it · §4.7.15 · cross-listed College 00 Reproducing the numbers proves the code makes the numbers. It does not prove the numbers mean what the paper says.
🍳 Philosophy of Science · the room where you cook it yourself

The Kitchen

Every other room on the Route asks whether to believe a paper. This one asks something smaller and harder to dodge: can you follow the recipe and make the dish in the photo?

“Can you follow the recipe and make the dish in the photo?”the only question this room asks
what the room is for

Cooked is not the same as good

A recipe can be followed exactly, produce exactly the dish in the photo, and still be a bad dish. The Kitchen does not ask whether the dish is any good. It asks whether the thing can be made at all, by somebody who wasn’t in the room when it was written.

This is not peer review. Peer review asks should we believe this. The Kitchen asks can you run this. An agent or a simulation built from a paper inherits everything in the paper — its assumptions, its data choices, its blind spots. Reproducing the paper’s numbers with the paper’s own code proves the code produces those numbers. It does not prove the numbers mean what the paper says they mean.

That distinction is the whole room. Every exhibit below is a case where somebody blurred it.

the map

Kitchen to research, term by term

In the kitchenIn the paper
The recipeMethods and code
The pantryThe data
The photo in the cookbookThe reported figure or number
“Season to taste”An unstated parameter — a knob the author knew about and left open
Grandma’s pinchTacit knowledge — a step the author didn’t know they were doing. The recipe is complete to the person who wrote it
“Your oven runs hot”Environment drift — software versions, hardware, a library that moved
The altitude adjustmentWorks in their kitchen, fails in Denver. Works on their data, fails on yours
Can you even get the ingredients?Is the data public at all?
The franchiseReplicability — see the second burner
Keep “season to taste” and grandma’s pinch apart. They look like cousins. One is a known variable left open; the other is an unknown step never written down. They fail differently and they get fixed differently — and neither of them is the third thing, which is the cook editing the recipe. That one gets its own line on the ticket.
two burners

Reproducibility and replicability are not the same stove

The field draws this line sharply, and the definitions below are quoted exactly from the National Academies’ 2019 consensus study — page 1 of the report itself, where the committee sets them out:

burner one · the corporate kitchen

Reproducibility

Reproducibility means computational reproducibility—obtaining consistent computational results using the same input data, computational steps, methods, code, and conditions of analysis.

Same recipe, same ingredients, same kitchen. Does it come out?

burner two · the franchise

Replicability

Replicability means obtaining consistent results across studies aimed at answering the same scientific question, each of which has obtained its own data.

Different kitchen, different supplier, a teenager on a Friday rush. Does it still come out?

The Academies add the part people skip: reproducibility should generally be expected; replicability is more nuanced, and in some cases a lack of it aids discovery. A franchise that can’t make the dish may have found something about the dish.

And the report names a third thing that is neither burner. Generalizability — “the extent that results of a study apply in other contexts or populations that differ from the original one.” That is the altitude adjustment as a formal concept: the dish came out, twice, and still won’t come out in Denver. The committee adds that “a single scientific study may entail one or more of these concepts” — which is why a ticket names which burner rather than passing or failing the paper.

Most automated tools only work burner one. Keep that in mind for the exhibit below.

two kinds of cooking

From the recipe, or rebuilt from the photo

you have the recipe

Cooking from the recipe

The paper ships code and data; you run them. Failure looks like an error message.

you only have the photo

Rebuilding from the photo

The paper ships prose and figures; you reconstruct a model from the reported numbers. Failure looks like a plausible answer that’s wrong.

Different skills, different failure modes, and a room that pretends they’re the same thing is lying to you. “No recipe published” is itself a finding, and it gets written on the ticket as one — not left blank.

exhibit i

The robot cook

Paper2Agent · Stanford · Nature 2026

Miao, Davis, Zhang, Pritchard & Zou built a system that converts a paper and its codebase into an AI agent: it reads the manuscript and the repository, builds an MCP server of callable tools, then generates and runs its own tests until they pass. The authors call the result a “virtual corresponding author.”

It is the cleanest published example of the Route’s front-door thesis — the paper stopped being a finished thing. The paper becomes something you query instead of something you read. And it is a robot cook working burner one.

What it did, at scale

  • 100 computational biology papers in, 74 agentified. 599 tools proposed, 593 passed automated validation.
  • 91.2 ± 1.6% on 300 benchmark questions (with Sonnet 4), against 80.3 ± 2.3% for a general agent pointed straight at the same repository.
  • The AlphaGenome case: 22 tools, all passing validation, in about 45 minutes for about US$14 on a laptop.

Why the failures are the interesting half

The paper names why the other 26 didn’t cook, and every reason is already on the map above:

What the paper reportsWhat the Kitchen calls it
missing executable codeno recipe
missing data or model artefactsempty pantry
environment or dependency failuresyour oven runs hot
non-generalizable scriptsthe altitude adjustment

And then the authors write the sentence this room was built to argue:

“the ease with which a paper can be transformed into an agent may itself serve as a practical measure of reproducibility.”Miao et al., Nature (2026), discussion

The taste test uses the cookbook’s own photo

The 91.2% is a burner-one number by construction, and the methods say so. The benchmark questions were derived from the repositories’ own tutorials, and the ground-truth answers were obtained by executing the original code and verifying against the tutorial outputs. Each tool is “validated against the reference codebase’s reported results and figures … and then locked to ensure reproducibility.”

That is a good, honest design for the question it asks. It is not a check on whether the paper is right, and this room exists because those two things get printed as one number.

the version trap, again

Same tool, two versions, different numbers

Paper2Agent exists as a preprint (arXiv:2509.06917v2, 16 October 2025) and as a Nature paper (received 13 October 2025, accepted 14 August 2026). Coverage quotes both without distinguishing them. They do not say the same thing:

ClaimPreprint v2Nature version
22 AlphaGenome tools took…about 3 hoursabout 45 min, US$14
What the “45 minutes” figure belongs to7 Scanpy tools7 Scanpy tools (US$13) and the 22 AlphaGenome tools
74 of 100 papers agentifiednot presentpresent
91.2% benchmark accuracynot presentpresent
A live one, logged: the copy of the Nature version used for this page still reads Published online: xx xx xxxx. It is a pre-final author copy, not the version of record. Pagination and final wording may move. Every number above was read out of the documents themselves, not from coverage — and the “45 minutes” that travelled through the press is the one figure that would have been wrong if it had been taken from the preprint.
exhibit ii · commit before the reveal

When the cook serves a different dish

Paper2Agent built an agent from AlphaGenome, DeepMind’s regulatory-genomics model. They asked it a question the AlphaGenome paper had already answered: which gene does a particular cholesterol-linked DNA variant actually act on?

The paper emphasised CELSR2 and PSRC1. The agent came back with SORT1. Paper2Agent reports this as a discrepancy and calls it a strength — the agent can re-evaluate published conclusions.

Before you read what the numbers say: when an agent built from a paper contradicts that paper, what have you got?
You said: a new claim. Then the bill is provenance. Before anything counts as new, somebody has to check whether it is — and the check has to run past the two papers in front of you, out into the literature behind them. Hold that thought and read the reveal.
You said: a bug. Defensible, and there is something wrong in the reasoning — see the arithmetic in the reveal. But be careful what you call a bug: the agent reached a defensible answer. A wrong route to a right destination is a real category, and it is not the same as broken code.
You said: a supplement. This is the strongest of the three, and the reveal will back you up: the paper was describing what the model predicts, the agent was ranking causal genes. Two questions. But you still owe the same provenance check as the first answer — because a supplement can also be a rerun of something already settled.
You said: something else. Good instinct. There is a fourth box and this room had to build it. Read on.

One variant, three papers, sixteen years

All three papers are standing on the same DNA base. Paper2Agent calls it chr1:109274968:G>T; AlphaGenome calls it rs12740374. Those are the same thing — dbSNP puts rs12740374 at chr1:109274968 on GRCh38, G>T. The same record gives its gene consequence: CELSR2 : 3 Prime UTR Variant. The variant sits inside one gene and, per Musunuru, acts on another about 150 kb away — which is a good part of why CELSR2 keeps getting named. (The position, the alleles and the consequence are all read off the dbSNP RefSNP record; the gene-distance reading is this page’s.)

The arithmetic the write-up glides past. Paper2Agent says the agent favoured SORT1 partly because of “a high quantile score (0.99983)”. Two paragraphs later it reports the competing genes’ scores: CELSR2 and PSRC1 at 0.99998 each. By AlphaGenome’s own measure, SORT1 ranks last of the three. The stated first reason for the agent’s pick points the other way.

GeneAlphaGenome quantile scoreGTEx liver eQTL
SORT1 — the agent’s pick0.99983 — lowest of the threeP = 1.1 × 10−65 — strongest
CELSR20.99998P = 4.7 × 10−46
PSRC10.99998P = 8.5 × 10−50

What actually carries SORT1 is the eQTL evidence and the biology — sortilin handles LDL and VLDL. Both of those come from outside the recipe. The agent didn’t out-cook the model; it went and got a different ingredient.

And the answer was already on the shelf

In 2010, Musunuru and colleagues published this, in Nature:

“a common noncoding polymorphism at the 1p13 locus, rs12740374, creates a C/EBP … transcription factor binding site and alters the hepatic expression of the SORT1 gene.”Musunuru et al., Nature 466, 714–719 (2010)

That paper did not stop at association. It knocked Sort1 down with siRNA and overexpressed it in mouse liver, and showed the gene moves plasma LDL cholesterol by changing hepatic VLDL secretion. Functional, causal, sixteen years old.

And AlphaGenome’s own figure caption says the variant “creates a CEBP binding motif” — Musunuru’s exact mechanism, sitting in the legend, pointed at the neighbouring genes.

The fourth box

“Rediscovery presented as re-evaluation.”the category this exhibit had to add

The agent is right. The paper it “corrected” wasn’t really wrong — it was answering a narrower question about what the model predicts. And the finding being celebrated as a re-evaluation of a published conclusion is the conclusion the field reached in 2010 with mouse livers.

None of this makes Paper2Agent a bad tool. The agent did something genuinely useful and did it in one prompt. The gap is in the paperwork around the result — and paperwork around a result is what this entire Route is about.

How far this claim goes, exactly — and it got further. Paper2Agent’s reference list contains no Musunuru 2010; it cites a sortilin review instead. AlphaGenome’s reference list at first would not extract from the PDF — the numbers came out without their text — so this page originally said “not found,” not “not cited.” It extracts now. Pulled page by page in layout mode, the list reads out complete, and the whole 31-page article was then searched in three extraction modes for “Musunuru,” the DOI nature09266, and the page range 714–719. Zero hits in all three. So: neither paper cites the 2010 work that settled this locus. The claim is now made, not hedged — and the hedge is left standing above it so you can see which one it was.
the form

The Kitchen ticket

Every paper that comes through the Route gets one of these. It is short on purpose — a ticket, not a review. Filled in below for the first exhibit.

Is there a recipe?
Yes — public repository, MIT licensed, with tutorials.
Is the pantry stocked?
Yes for the demonstration cases; the agent pipeline needs each paper’s own data, which is where 26 of 100 fell over.
Any “season to taste”?
Model choice. The headline accuracy is Sonnet 4; the baseline was also run on a newer model and moved.
Any grandma’s pinch?
Not identified. Would need an outside run to find out — which is the point of the room.
Recipe audit & adjustment
Yes, by design. The system finds and repairs what it inherits — broken dependencies, dead file paths, typos, retired APIs — then reports that the dish came out. A repaired recipe producing the photo is not the original recipe producing the photo. The audit is what was found wrong; the adjustment is what got changed. Both belong on the ticket, and neither is “season to taste” or grandma’s pinch.
Did it come out?
Reproduced — on the tool level, against the codebase’s own outputs.
Which burner?
Burner one. No new data was collected anywhere in it.
Cooked from the recipe, or rebuilt from the photo?
From the recipe — it only works on papers that ship one.

Next ticket up: the aerial-fiber entanglement paper behind QuantumPulse — no recipe published, rebuilt from the photo, one burner, and the photo itself changed between preprint and publication. A different failure mode in every row.

where you are

The Kitchen’s place on the Route

A claim gets caught in the weather, then it becomes a paper, and then two rooms open at once. The Docket asks whether there is enough evidence to know where this belongs. The Kitchen asks whether you can run it. Those are independent questions — a paper with no code can fail here and still have a strong case on the Docket.

caught in the wild → THE PARK BENCH ↓ preprint / conference paper ↓ ↓ THE DOCKET THE KITCHEN ← in parallel is there a case? can you run it? ↓ ↓ → survives both ← ↓ THE SLEDGEHAMMER WING every swing goes here first ↓ the verdict routes it ↓ ↓ it connected it was the ruler (stays: a paradigm) → THE CALIPER ROOM ↓ ↓ → nothing new, nobody arguing ← ↓ THE TEACHERS’ LOUNGE no clock on this duct tape → pushpin → plexiglass → frame

The far end has no clock on it. A case can sit in the Sledgehammer Wing or the Caliper Room for a week or for thirty years. It moves to the Teachers’ Lounge when nothing new turns up and nobody is arguing any more — and even that is not the end, because the Lounge has its own ladder: duct tape → pushpin → plexiglass → frame. How a case is pinned tells you how settled it is.

Nobody grades a meal that hasn’t been cooked, which is why the Kitchen sits before the two verdict wings. And the ticket travels with the exhibit. If a claimed paradigm break can’t be re-run at all, that is evidence pointing toward the caliper reading — the Kitchen doesn’t grade, but its ticket is an input to whoever does.

sources

What’s behind the numbers

  1. Miao, J., Davis, J. R., Zhang, Y., Pritchard, J. K. & Zou, J. “Reimagining research papers as interactive and reliable AI agents.” Nature (2026). doi:10.1038/s41586-026-11044-y — open access. Received 13 Oct 2025, accepted 14 Aug 2026. All Paper2Agent figures on this page are from this version. The copy read here is a pre-final author copy (“Published online: xx xx xxxx”).
  2. Paper2Agent preprint. arXiv:2509.06917v2, 16 Oct 2025. Used only for the version-trap comparison. Code: github.com/jmiao24/Paper2Agent (MIT).
  3. Avsec, Ž. et al. “Advancing regulatory variant effect prediction with AlphaGenome.” Nature 649, 1206–1218 (issue of 29 January 2026; published online 28 January 2026; received 16 May 2025, accepted 4 December 2025). doi:10.1038/s41586-025-10014-0 — the CELSR2 / PSRC1 figure caption and the C/EBP motif observation. Volume and pages read off the article’s own running heads.
  4. Musunuru, K., Strong, A., Frank-Kamenetsky, M. et al. “From noncoding variant to phenotype via SORT1 at the 1p13 cholesterol locus.” Nature 466, 714–719 (5 Aug 2010). doi:10.1038/nature09266 — abstract read; full text paywalled.
  5. National Academies of Sciences, Engineering, and Medicine. Reproducibility and Replicability in Science. Washington, DC: The National Academies Press, 2019. doi:10.17226/25303, ISBN 978-0-309-48616-3, 257 pp. Definitions and the generalizability passage read from page 1 of the report; they also appear in the 4-page Consensus Study Report Highlights (May 2019). Report read here as the free NCBI Bookshelf edition, NBK547537.
  6. dbSNP RefSNP reportrs12740374. Position chr1:109274968 (GRCh38.p14), alleles G>T, gene consequence CELSR2 : 3 Prime UTR Variant (transcript NM_001408.3:c.*919). Read from the RefSNP record itself.
What’s real / what’s mine. ● Real: every number, quotation, date and gene on this page, read out of the documents above rather than out of coverage. ◐ Mine: the kitchen metaphor, the ticket, the “rediscovery presented as re-evaluation” category, and the reading of the SORT1 case. Those are argument, not findings — and the arithmetic they rest on is printed above so you can check it without taking my word for anything.