An artificial intelligence system trained to mimic how human pathologists visually search tissue slides for cancer outperformed other general-purpose AI models at the task, according to a study published in July in the journal Nature. The approach reduced missed cancers but came with a higher false-positive rate, a trade-off researchers say could still make it useful as a screening aid for physicians.
The method addresses a structural mismatch in how many AI systems analyze pathology slides. Standard algorithms often examine preselected regions or divide a slide into fixed-size patches. A human pathologist works differently, panning across the tissue, zooming in and out, and pausing over areas that raise suspicion. A whole slide can contain billions of pixels, while evidence of cancer may occupy only a tiny patch.
"You don't start by inspecting one square meter of ground," said study co-author Zhi Huang, an assistant professor of pathology and laboratory medicine at the University of Pennsylvania. "You scan the landscape first and then swoop in for a closer look."
Capturing search behavior as training data
The researchers called their training approach Pathology-CoT, short for "chain of thought." Instead of learning only from final labeled images or diagnoses, the AI learned from the observable actions pathologists take during their search - where they move, where they linger, and when they change magnification.
Eight pathologists used a custom tool that recorded their movements across slides. The raw logs were noisy: a pathologist might drift, overshoot a region, or fiddle with magnification. The team filtered out incidental movements and retained moments of deliberate attention, such as sustained panning or lingering over one view. They checked those regions against eye-tracking data to confirm the software captured where pathologists were actually looking.
For each inspected region, a vision language model drafted a short rationale explaining what features were visible and why the area warranted examination. Human pathologists could accept, edit, or reject the rationale, creating additional training data. In one example, the AI flagged a portion of a slide as potentially metastatic and suggested zooming in to look for atypical cells.
The resulting tool, called Pathology-o3, scans a slide at low resolution, uses a model trained on pathologists' behavior to choose regions worth a closer look, then sends higher-resolution views of those regions to a vision language model for analysis.
Testing against general-purpose AI
Huang said the goal was not to beat specialized models built for specific cancer types. Instead, the team wanted to see whether their training approach could help a general-purpose AI navigate pathology slides more effectively. They compared Pathology-o3 to other general-use systems, including OpenAI's o3, using lymph node tissue slides from colorectal cancer cases that human pathologists had already labeled.
Pathology-o3 correctly identified slides positive for cancer 100% of the time. However, 15.5% of the slides it flagged as positive were actually negative. OpenAI's o3 correctly identified positive slides 87.5% of the time, but 53.3% of its flagged positives were false alarms. The researchers designed Pathology-o3 to err on the side of flagging something for another look rather than potentially missing cancer, which helps explain the false-positive rate, Huang said.
On an independent dataset the algorithms had not seen before, Pathology-o3 correctly identified positive slides 97.6% of the time, with a false-positive rate of 37.1%. Mohammad Asadi, a data scientist at Stanford University who was not involved in the research, said this result suggests the system can work with slides from a different source, though it does not yet show that using the tool makes pathologists more accurate or efficient in practice.
How the tool might be used
Asadi noted that Pathology-o3 is not precise enough to diagnose patients on its own. It could still serve as a prescreening tool that points a human toward specific regions worth double-checking. Showing a pathologist a specific area to inspect, rather than declaring an entire slide suspicious, may be an advantage.
"The right question isn't whether it beats a pathologist," Huang said. "It's whether a pathologist working with it catches more [cancer cases] and works faster." The current study did not address that question, but the team's next experiment will test pathologists on the same cases with and without Pathology-o3, measuring both detection rates and time spent.
Asadi wants an even tougher test: trials conducted across multiple hospitals that measure accuracy, speed, and pathologists' workloads. He said such trials should assess the burden of false alarms and whether doctors recognize when the AI is wrong. Cancer diagnoses can also require information from multiple slides, stains, and a patient's medical history, while the current system reads one slide at a time.
The researchers applied their training approach to several existing vision language models and found that performance consistently improved afterward. For Huang, that is the central finding. "The takeaway isn't our system," he said. "It's that the missing ingredient has been sitting in hospitals this whole time."
Why this matters for science and research professionals
The study demonstrates a transferable training methodology rather than a one-off model. The core insight - that recording and codifying expert search behavior can improve AI performance - applies beyond pathology to any domain where professionals visually inspect complex data, from radiology to materials science. For researchers working in AI for Science & Research, the approach offers a template for building tools that complement human expertise instead of attempting to replace it. The method also highlights how domain-specific behavioral data, often already generated in clinical and lab workflows, can serve as a training signal that general-purpose models lack. Professionals developing or evaluating AI for AI for Healthcare applications should note that the study's most important benchmark is not standalone accuracy but whether the human-AI pairing catches more cases and reduces workload - a standard that will likely shape procurement and validation requirements in regulated settings.
Your membership also unlocks: