Connect with us

NEWS

The Unused Search Traces That Taught AI to Find Cancer

Penn researchers trained cancer AI on pathologists’ unused search logs, hitting 100% recall on Stanford slides while leaving extra flags for those same doctors.

Published

on

Pathology-o3, a Penn-built agent, reached 100% recall for metastatic cancer on Stanford lymph-node slides after training on 10.6 hours of pathologist search. The system learned from eight doctors’ pans, zooms, and pauses in a routine viewer, then released a public set of 5,222 reasoning rounds. Zhi Huang, an assistant professor of pathology and laboratory medicine at the University of Pennsylvania, said that viewing trail had been sitting in hospitals unused.

The July study in Nature Biomedical Engineering put that behavior-guided agent against general vision models, including OpenAI’s o3, on colorectal cancer lymph nodes. It was built to over-flag rather than miss tumor, which is the bargain those same pathologists would take on in a real lab.

Eight Pathologists, 10.6 Hours, One Unused File

Most pathology AIs still train on what a doctor leaves at the end of a case: a circled tumor or a signed diagnosis. Huang’s group wanted the middle of the hunt. Diagnosing a gigapixel slide is a moving search, closer to panning a map than grading a cropped photo, and the evidence of cancer can occupy only a small field.

Huang compared the job to a search-and-rescue helicopter. “You don’t start by inspecting one square meter of ground,” he said. The team recorded eight pathologists at Stanford Medicine (four attendings, two fellows, and two residents) as they worked 137 lymph-node slides from 25 colorectal cancer cases on the nuclei.io whole-slide viewer used in the lab. The logs cover 921 sessions and 10.6 hours of inspection. The Perelman School of Medicine IRB listed the protocol as #857258, and the slides were scanned on a Leica Aperio at 40x (0.25 µm per pixel).

Huang announced the journal paper on 24 July 2026, the day it went online. Penn Pathology and Laboratory Medicine passed the post along three days later.

THE TRACE THAT TRAINED THE AGENT

  • The people: Eight Stanford pathologists, plus two outside reviewers at UCSF and Penn.
  • The hours: 10.6 hours of diagnostic sessions across 137 slides and 25 cases.
  • The text: 5,222 conversation rounds, with about 152 words on each low-power inspect and 82 on each 40x peek.
  • The noise: More than 257 viewport events per slide, logged at about 10 Hz.

Raw, that stream is almost unusable. Feeding every viewport to a vision model would produce over 500,000 visual tokens per case, past the context window and full of fidgets, overshoots, and magnification tweaks. Hospitals already store versions of these logs. Almost none of that exhaust has been turned into training data.

The Recorder Keeps the Pauses and Drops the Fidgets

The group’s answer is a piece of software they call the AI Session Recorder. It sits on a standard whole-slide viewer, reads the timestamped field of view, and cuts the continuous mouse trail into microscope-like commands. Charles Herndon, a pathologist at the University of California, San Francisco, and Shunsuke Koga, a pathologist at the Hospital of the University of Pennsylvania, then checked the drafted reasons.

HOW THE RECORDER CUTS THE NOISE

  • Inspect: A low- or medium-power look at 5x or 10x, kept when the viewport stays still more than one second or pans more than two seconds.
  • Peek: A fast 40x look at cells, paired with a bounding box the model can crop.
  • Drafted reason: A vision model writes why the region was worth a look and what is in the field; a pathologist accepts, edits, or marks “no zoom here.”
  • Timing: One reviewer averaged 18.8 seconds to check a drafted round, against 106.2 seconds to type the same work from scratch; a second reviewer averaged 16.3 seconds against 85.2 seconds of typing. The team called that about six times faster than typing and three to four times faster than dictation, and more than 80% of drafts needed no edits.

Sheng Wang and Ruiming Wu, the co-first authors at Penn, then used those cleaned traces to train Pathology-o3. A YOLOv8 behavior predictor proposes regions. A vision-language model reads the crops and writes a step-by-step case note. First-author credit also lists Jeanne Shen, a Stanford pathologist who helped gather the Stanford cases. The public release packages the 5,222 expert reasoning rounds on GitHub with the recorder code and a demo agent.

Pathology-o3 Caught Every Positive Slide at Stanford

The internal test used gastrointestinal lymph-node slides from Stanford Medicine, some already labeled for metastatic colorectal cancer. Huang said the point was not to beat a disease-specific detector trained on one cancer. It was to see whether a general vision model could search a slide more like a person once it had been shown where people look.

STANFORD LYMPH-NODE RESULTS

System Precision Recall Accuracy
Pathology-o3 84.5% 100% 75.4%
OpenAI o3 46.7% 87.5% 57.8%

OpenAI’s o3 was the next-best general model in that lineup. Models that do not think with a zoom trail split in ugly ways: Phi-4 reached 91.7% recall but only 42.3% precision, and Llama-4 reached 75.0% precision while recalling only 6.2% of the positives. Across the rest of the baselines, precision ran from 33.3% to 65.6%, recall from 18.2% to 91.7%, and accuracy from 42.2% to 65.6%. Huang’s group designed Pathology-o3 to flag a region for another look rather than risk a miss, which is why some negative slides still come back as positive.

Swedish Scanners Did Not Break the Search Policy

The behavior predictor was trained on Stanford viewing habits and then run, with no extra tuning, on 321 slides from Sweden. That set, drawn from the Swedish LNCO2 lymph-node collection, uses different scanners and magnifications (Aperio at 20x and Hamamatsu at 40x in the full register of 50 consecutive colon-adenocarcinoma cases). Pathology-o3 still reached 97.6% recall, with 62.9% precision and 69.4% accuracy. Its precision was more than double Gemini’s 23.5% on the same external test.

Qualitative overlays in the paper show the predictor landing on tinctorial shifts, broken architecture, and capsule edges that a senior attending also worked. When the group let unguided vision models pick their own fields, those boxes were scattered and low-yield. Completeness (the share of expert-viewed regions recovered) and efficiency (the share of picked regions that were clinically relevant) both fell. The analytical skill in these models was trained on static crops. The search skill was not.

Order of the regions barely mattered. Forward, reverse, and random presentations moved the scores only a little, which fits a hunt-for-tumor task. Capping how many regions the agent could inspect, a stand-in for chain-of-thought length, lifted recall as the list got longer while precision stayed high. The paper is frank that this may not transfer to jobs where the path itself is the diagnosis, such as grading prostate architecture or counting mitoses along a tumor edge.

Who Pays for the Extra Flags?

Mohammad Asadi, a Stanford data scientist who was not on the paper, said the error mix is not tight enough for the model to diagnose patients alone. It can still be useful if it points a person at a field rather than branding a whole slide as suspicious. That is also the catch for the people whose search traces trained it. A region-level flag can be checked. A slide-level alarm is just more queue.

Asadi wants a harder trial than a held-out slide set: several hospitals, measuring accuracy and speed, plus the extra work from false alarms and whether doctors notice when the model is wrong. The present system reads one slide at a time. Real cancer sign-out often needs several slides, extra stains, and the rest of the chart.

WHAT WE KNOW

  • The scores: Perfect recall on the Stanford internal set, 97.6% recall on 321 Swedish slides, with precision that drops when the scanner changes.
  • The design: The agent is tuned to over-call, and it was not compared head-to-head with unaided pathologists.
  • The scope: One slide, one stain, one search-and-find task, colorectal lymph-node metastasis.

WHAT IS UNCONFIRMED

  • Lab speed: No published test yet shows that a pathologist with Pathology-o3 is faster or catches more tumor.
  • Alarm load: No multi-hospital measure of how many extra fields a doctor must dismiss, or how often those flags are believed when they are wrong.
  • Clinic status: No FDA authorization for this agent; it is a research system, unlike a small set of whole-slide tools already cleared for other tasks.

A separate risk sits inside the labeling loop. Because a model drafts the “why” text before a person edits it, the expert can be pulled toward the machine’s wording. The authors flag that anchoring bias themselves. Viewer logs also differ by software, so each lab still needs an engineering pass before collection becomes passive.

Ordinary Vision Models Improved After a Dose of Search History

The result Huang cares about most is not Pathology-o3 as a product. When the same viewing policy was attached to several existing vision-language backbones, precision rose by an average of 11.8% and recall by 17.9% against a thumbnail-only baseline. Accuracy rose 9.2% with the learned policy. Guidance from a senior attending gastrointestinal pathologist, a practical ceiling, raised precision 11.9%, recall 19.1%, and accuracy 7.1%. The compact predictor closed most of that gap, which is how a short recording campaign becomes a reusable search prior instead of a forever annotation project.

The paper puts that in contrast with agents that let a general model decide where to zoom from textbooks and the open web. Pathologists’ search habits are tacit. They are not written down, and unguided models wander. Multiple-instance methods that tile an entire slide were left out of the comparison on purpose, because they do not follow a human search and cannot be trained on these traces.

A first preprint went up on 6 October 2025 (v2 on 13 October). The journal accepted the manuscript on 8 June 2026 and published it on 24 July. Startup funds from the Perelman School of Medicine supported the work. Code and the chain-of-thought set are public; the Swedish and a later skin-cohort slides sit behind application-only archives.

The Lab’s Next Trial Puts Pathologists Back in the Loop

The study never claimed to outrun a pathologist. Huang has already described the experiment he wants next: the same cases, read with and without Pathology-o3, scoring what is caught and how long it takes. Until that runs, the honest product is a prescreen that shows a field, not a signature on a report.

The takeaway isn’t our system. It’s that the missing ingredient has been sitting in hospitals this whole time.

Zhi Huang, assistant professor of pathology, University of Pennsylvania

“The right question isn’t whether it beats a pathologist,” Huang said. “It’s whether a pathologist working with it catches more [cancer cases] and works faster.” He also drew a hard line on autonomy. “I wouldn’t claim it should diagnose on its own.”

Disclaimer: This article is news reporting on a peer-reviewed research study and is for information only. It is not medical advice, a diagnosis, a treatment recommendation, or an assessment of any patient’s slides. Readers should consult a board-certified pathologist or other qualified physician about clinical decisions, and laboratory directors should consult regulatory counsel before considering any research model for patient care. Figures, study status, and tool capabilities reflect the cited papers and datasets as published and may change with later trials, software versions, or regulatory review.

Harry is the editor of BROAD BROWSE, which he owns, runs and largely writes himself as an independent publication. The site is deliberately wide, and keeping ten sections accurate with one editor depends on a rule he has followed through a decade in journalism, from reporter to editor: every section has its own primary record, and the article starts there. For business that means the filing and the earnings call transcript, for science the paper and its underlying data, for sports the official result, for auto and technology the product in his hands, for news the statement or the court document. Entertainment, lifestyle, travel and gaming get the same treatment, with the release, the itinerary or the game itself checked before writing begins. Readers come from many countries, so figures are given with context and checked before they are published. Corrections are made on the article with a dated note, and the site's corrections policy is public. He answers reader mail personally at support@broadbrowse.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending