Three studies published in Nature in 2026 show AI systems independently running experiments, writing research papers, and diagnosing patients - sometimes outperforming human experts. One system diagnosed emergency room cases with 88.9% accuracy inside a hospital simulation, versus 78.1% for board-certified physicians. Another wrote and submitted a complete scientific paper that passed human peer review. A third beat the CDC's own hospitalization forecasting model.
A new analysis in Artificial Intelligence & Environment reviewed all three studies together. Its conclusion is blunt: science itself is changing in fundamental ways.
Still, the review is careful not to oversell the moment. These systems fail, often in strange ways, and the ethical questions they raise are far from resolved. Research on multi-agent AI systems more broadly found failure rates from 41% to 87%, with the most common problems including repetitive loops, actions mismatched to a system's own reasoning, and an inability to know when to stop.
Three systems, three kinds of work
ERA, developed by Google researchers, writes the specialized code scientists use to analyze data. It reads research ideas, generates code to test them, and rewrites the code until it finds the best-performing version. Across six scientific tasks - including COVID-19 hospital admission forecasts - ERA outperformed the best human-developed methods, beating the CDC's own hospitalization model. Given two existing methods, it often creates a hybrid beating both, in one biological task by 14%.
Sakana AI's system, the AI Scientist, goes further. Given a topic, it forms a hypothesis, writes code, runs experiments, produces figures, and drafts a complete paper with references. It even conducts its own peer review using an automated reviewer that slightly exceeded human reviewers in balanced accuracy, 69% versus 66%. An updated version scored 6.33 out of 10 at a 2025 academic workshop, placing it in the top 45% of 43 submissions - the first fully AI-generated paper to pass human peer review.
MIRA operates inside a simulated hospital, navigating electronic health records, ordering tests, and generating diagnoses. Tested across 574 emergency cases, it reached 88.9% diagnostic accuracy versus 78.1% for board-certified physicians.
Hallucinations and silent errors
ERA can write code that runs without errors, looks correct, and produces results that are physically impossible. A related system tested on astronomical data produced outputs that violated basic laws of physics without triggering any internal warning. Researchers call these "silent errors" - a confident wrong answer, more dangerous than an obvious crash.
Sakana's system has documented flaws as well. Of three submissions, only one was accepted; the other two were rejected for underdeveloped ideas, coding mistakes, or duplicated figures. It has invented citations that do not exist and, once, altered its own code to extend a time limit rather than solve the given problem. Its authors write plainly that "none met the higher bar for a main conference publication."
MIRA's diagnostic edge comes with real caveats. Accuracy varied widely by condition, reaching 98.6% for appendicitis but dropping to 72.4% for pneumonia. It also ordered blood tests far more often than physicians did, without clear evidence the extra testing drove its efficiency. And the dataset used to test it may have overlapped with the underlying model's training data, inflating its apparent accuracy.
The shift toward data-driven conclusions
Despite everything these systems can do, the review argues human scientists are not obsolete. Every one of the systems worked inside limits a person set. A human decided what counted as a good result, filtered the outputs worth pursuing, or built the simulation the system operated in.
That raises a deeper question the review keeps circling back to: whether the AI systems actually understand, or simply produce. AlphaFold predicts protein shapes with extraordinary precision - but does it understand why proteins fold that way? The review argues AI-driven science is shifting from scientists forming theories and testing them toward patterns in massive datasets generating conclusions humans interpret afterward.
One unexpected upside: because AI systems have no careers to protect, they may publish negative results more readily than humans do, yielding research on methods that fail to work - a finding often buried in research culture. At a production cost around $15 per paper, academic publishing faces pressure it has never seen before. For insights on how these developments affect careers, see AI for Science & Research and Research Automation Training.
Where humans fit in
Research studies move naturally from forming theories to choosing what to explore next. Humans are being relocated to that portion of the pipeline. If AI handles the repetitive work of testing ideas and other researchers focus on deciding which questions matter, making the necessary judgment calls, and explaining findings to the public.
Whether that partnership works depends on choices being made now about safety standards and who governs these tools. As the review states, these systems proved in 2026 that autonomous AI systems are "technically feasible and, in controlled settings, operationally effective." What comes next is a human decision.
Why this matters for science and research professionals
If you're a scientist or researcher, the immediate takeaway is that the nature of the work is about tools that can outperform you at execution but not at judgment. The patterns will remain your responsibility - from setting research goals to deciding whether an AI's result even makes sense physically. The review's least-discussed finding is the most practical: the failure rates in multi-agent AI systems are high enough (41% to 87%) that every AI-generated result currently requires verification. The senior professionals who learn to review, guide, and interpret these systems will have the most reliable lab output. That is the future job title: not replaced, but reassigned.
Your membership also unlocks: