Tool Turns 74 of 100 Biology Papers Into AI Agents
Stanford's Paper2Agent turns a paper and its code into a tool a chat AI can run; it worked on 74 of 100 tested biology papers.
Estimated reading time: 6 minutes
TL;DR
- A Stanford team has published Paper2Agent, a system that turns a research paper and its code into an AI assistant that can run the paper’s methods on request in plain English. The study appeared in Nature on September 16.
- In a trial on 100 computational biology papers, it produced working agents for 74, IEEE Spectrum reported. The rest failed, often because of incomplete code, missing documentation or software that could not be made to work.
- In the team’s 2025 preprint, an agent built from the AlphaGenome paper answered all 30 test questions correctly. Two other AI tools answered 21 and 15.
- By connecting agents built from different papers, and with researchers guiding the analysis, the team produced two genetic leads, for ADHD and psoriasis. Both are computational hypotheses, not laboratory findings.
- The team also tried the system on statistics, econometrics and astrophysics papers, but its published demonstrations focus on biology papers that come with code.
What happened
Reusing a scientific paper’s methods usually means reading the text, then working through someone else’s code to rebuild the analysis. A Stanford team led by Jiacheng Miao and James Zou built Paper2Agent to take over that step, according to Stanford Medicine. The study was published in Nature on September 16, 2026.
The system reads a paper and its code repository and uses several AI agents to package the paper’s methods as tools. It then tests those tools against the paper’s reference results, tries to fix the ones that fail and drops those it cannot repair, IEEE Spectrum reported. The tools are served through the Model Context Protocol (MCP), a standard connector that lets an AI assistant call another program’s functions.
Connected to an assistant such as Claude Code, Anthropic’s AI coding tool, the tools let a user ask a question in plain English and have the assistant run the paper’s own code to answer it. Stanford Medicine’s account adds that the human authors can still supply context the paper leaves out, by answering the agent’s questions.
For AlphaGenome, a model that predicts how DNA mutations affect gene activity, building the agent took about 45 minutes on a personal laptop, with no human intervention, and cost less than $15 in computing, IEEE Spectrum reported. It produced 22 tools, all of which passed the system’s automated checks. The team’s 2025 preprint had put the build time at about three hours.
In that preprint, the team compared the AlphaGenome agent with two alternatives: Claude Code given direct access to AlphaGenome’s code, and Biomni, a specialist biomedical AI tool. On 15 questions based on AlphaGenome’s tutorials, the agent answered all 15 correctly, against 9 for Claude Code and 6 for Biomni. On 15 further questions the team wrote itself, using variants and cell types absent from the original materials, the agent again scored 15, against 12 and 9. The team checked each answer by hand against results from running the original code.
At larger scale, the team ran Paper2Agent on 100 computational biology papers and turned 74 into working agents, IEEE Spectrum reported. The other 26 failed, often because of incomplete code, missing documentation or software that could not be made to work. The magazine also reported that the team tried the system on papers in statistics, econometrics and astrophysics, though its demonstrations focus on biology.
The team also linked agents built from different papers into what it calls an AI co-scientist. In the ADHD case, described in the preprint and by Stanford Medicine, an agent built from the AlphaGenome paper worked with one built from a genetic study of ADHD. The researchers chose the two papers and steered the co-scientist toward one of the hypotheses it proposed. Out of 209 candidate variants, it singled out one, rs1626703, which it predicts changes how the gene MPHOSPH9 is spliced in a type of neuron. Zou told Stanford Medicine the link had not been reported before.
In a second case, IEEE Spectrum reported that three agents examining psoriasis genetics pointed to a little-studied gene, GPR137, as a likely cause and proposed 10 ways to test the idea. A researcher picked one. The analysis found that silencing GPR137 caused changes in gene activity similar to those caused by the psoriasis-linked variant in immune cells.
Paper2Agent’s code is open source under an MIT license on GitHub, with ready-made agents for the AlphaGenome, ScanPy and TISSUE papers. The team had built more than 100 paper agents by the time of publication and hopes most papers will eventually have one, Stanford Medicine reports.
What this means (and what it does not)
If it works as described, Paper2Agent could let researchers apply a published method directly instead of rebuilding it from code. Zou told Stanford Medicine that knowledge has long been stored as “very passive artifacts,” and he presents paper agents as an active alternative.
All performance figures in this article come from the research team. The team wrote the test questions and graded the answers by hand, a limitation its preprint acknowledges. Stanford Medicine’s account comes from the university’s own news office.
The results do not show that paper agents can replace expert judgment or laboratory work. The ADHD and psoriasis leads are computational hypotheses, and researchers chose the papers, steered the analysis and picked the follow-up test. Zou told Stanford Medicine that collaboration between agents should be closely guided and monitored. The 74-of-100 conversion rate applies to computational biology papers that already came with code.
What we still do not know
- How much human help the build-and-test process needed across the 100-paper trial is not reported. For AlphaGenome, the team reports none.
- Whether the approach works on papers without code, such as theoretical or survey papers, is not addressed in the sources.
- How paper agents perform on questions and data chosen by people outside the research team has not been reported.
- The 26 failed conversions are described by their most common causes, with no count of how often each occurred.
- Neither the MPHOSPH9 variant nor the GPR137 lead is reported as confirmed by independent laboratory experiments.
- The benchmark and build-time figures in the 2025 preprint may differ from those in the peer-reviewed version. A figure reported elsewhere, 91.2% accuracy against 80.3% for a Claude Code baseline on 300 questions, appears to come from the Nature paper. That paper is behind a paywall, so the figure could not be checked against it and is not used here.
Sources & Bylines
Every source cited in this article, gathered in one place.
- https://www.nature.com/articles/s41586-026-11044-y
- https://spectrum.ieee.org/paper2agent-ai-agents-research-papers — Elie Dolgin
- https://arxiv.org/abs/2509.06917
- https://med.stanford.edu/news/all-news/2026/09/ai-agents-talk.html — Hanae Armitage
- https://arxiv.org/html/2509.06917v2
- https://github.com/jmiao24/Paper2Agent
Editorial check, counted automatically
- 6 sources cited
- 18 inline-linked claims
- 0 unsourced claims found
- 0 banned words found
- 1 numbers without context
Also available in Portugues (BR)