Diagnosis UPSET Stuns Medical Staff

Healthcare professional uses tablet with medical holograms
Photo: metamorworks / Shutterstock

Artificial intelligence is already beating human doctors in some medical tests, and the latest study says it did so with the same patient records the doctors had.

Quick Take

  • An OpenAI reasoning model outperformed two experienced doctors in a Harvard and Beth Israel Deaconess Medical Center study.
  • The model worked from the same electronic health record information the physicians had at the time.
  • Broader research still shows mixed results, with AI doing well in some tests but falling behind expert physicians overall.
  • The split helps explain why health care leaders see promise, but not proof, that AI is ready to replace doctors.

What the Harvard Study Found

Researchers at Harvard Medical School and Beth Israel Deaconess Medical Center tested an OpenAI reasoning model against physicians on patient diagnosis and care decisions. NPR reported that the model matched and often outperformed the doctors when both sides were limited to the same electronic health records and other information already in the chart. The result matters because it was not based on extra clues the AI alone received.

The study adds to a growing set of headline results showing that AI can do well on narrow medical tasks. Science News said the model was more likely than physicians to include the correct diagnosis, or something very close to it, in its answers. Nature reported that the AI got the diagnosis correct or almost correct in 67 percent of cases, compared with about 50 to 55 percent for the two doctors in the experiment.

Why the Result Is Getting So Much Attention

The new finding hits a nerve because it goes straight to one of the most basic parts of medical care: figuring out what is wrong. Patients often wait weeks for appointments, face rushed visits, and leave without clear answers. A tool that can score better than trained doctors on written cases raises a simple question about who, or what, gets trusted first when the stakes are high. That is why the result spread so quickly beyond medical journals.

The result also fits a wider public frustration with systems that feel slow, expensive, and overburdened. When a machine can outperform humans on a controlled medical test, people on both the left and the right tend to ask whether the old system is serving patients or protecting habits. The study does not prove AI is ready for every bedside decision, but it does show that the technology is no longer a toy in diagnostic work.

What the Wider Research Says

Broader evidence is less dramatic than the newest headline. A meta-analysis of 83 studies found an overall diagnostic accuracy of 52.1 percent for generative AI and no significant difference from physicians overall, but it also found AI performed significantly worse than expert physicians. That gap matters because many public claims about AI use average results, while the strongest results often come from narrow tasks with fixed answers and clean scoring rules.

Earlier reviews reached a similar middle ground. One review found AI was comparable to clinicians in some settings and better than less experienced doctors, especially in image-based diagnosis. Another study found a chatbot on its own outperformed doctors who only had internet search and medical references, but doctors who used their own large language model kept up with the chatbots. The pattern suggests AI can help, but it does not yet settle the question of who should lead care.

What Comes Next for Hospitals and Patients

The practical issue now is not whether AI can win a benchmark. It is whether hospitals can use it safely in real care without creating new errors, delays, or blind spots. The Harvard study points to strong potential in triage and diagnosis, but the broader literature still warns against treating one good result as a full answer. For now, the evidence says AI can beat doctors in some tests, yet expert-level medicine still belongs in the real world, not just the lab.

That is why the strongest reading of the research is neither hype nor panic. AI is proving that it can spot patterns and make medical judgments faster than people in selected settings. At the same time, the larger body of evidence shows that performance varies by task, by model, and by the skill of the clinician being compared. In medicine, that difference is the line between a useful tool and an unsafe shortcut.

Sources:

reason.com, npr.org, pmc.ncbi.nlm.nih.gov, sciencenews.org, teledirectmd.com, news.creeta.com, sciencedaily.com, nytimes.com, erictopol.substack.com