Last updated: May 7, 2026
TLDR: A new study from Harvard Medical School and Beth Israel Deaconess Medical Center, published in Science on May 1, 2026, found OpenAI’s o1 model gave the right or near-right diagnosis on 67% of 76 real ER cases. The two human physicians scored 55% and 50%. Lead authors: Arjun Manrai and Adam Rodman. Importantly, the human comparators were internal-medicine attendings, not ER specialists — a nuance the headlines have mostly skipped. The authors are clear: this does not show AI is ready to make life-or-death calls alone. It shows AI could be a strong second opinion.
Headlines will say “AI beat the doctors.” The study says something more interesting and more useful. Here’s the actual experiment, the one example everyone will quote, and the case for AI as a colleague rather than a replacement.
What the Harvard study actually found
The researchers pulled 76 real patients who walked into the Beth Israel emergency room. For each patient, they captured the same electronic medical record information that was available at the time of triage. Two internal-medicine attending physicians wrote diagnoses for each case. OpenAI’s o1 and 4o models were given the exact same records (no preprocessing) and asked to do the same thing.
Two more attending physicians then graded all the answers without knowing which came from a human and which came from an AI. The o1 model produced a correct or close-to-correct diagnosis 67% of the time. Physician A scored 55%. Physician B scored 50%. The gap was widest at first triage, when the available information was thinnest, which is the moment in the ER where mistakes are most expensive and where a strong second opinion would help most.
The lupus catch — one example everyone will quote
One of the cases involved a patient who kept getting worse without an obvious explanation. The o1 model flagged a possible history of lupus that the human physicians had not surfaced. Adding lupus to the differential diagnosis ended up explaining the trajectory of the patient’s symptoms.
That’s exactly the kind of catch large language models should be good at. They’ve read millions of medical case summaries; they don’t get tired; they don’t anchor on the first plausible explanation. They also don’t have malpractice insurance or a license, which is why “second opinion that flags rare conditions” is the role the authors actually defend.
Why AI beat ER doctors is a misleading headline
The two physicians whose diagnoses got compared to o1 were internal-medicine attendings, not ER specialists. ER physician Dr. Kristen Panthagani pointed this out in public commentary on the study, and it’s a fair criticism. ER doctors are trained specifically for the early-triage decisions the study tested; the comparison is less impressive when the AI is being benchmarked against doctors whose specialty is something else.
That doesn’t erase the result. It just changes its framing. What the study actually shows is that o1 is a strong diagnostic reasoner on the kinds of cases that show up at first triage, and that it outperformed two doctors who, while board-certified physicians, weren’t ER specialists. “Strong reasoner that catches things tired humans miss” is genuinely useful. “Replaces ER doctors” is not what the data says.
What the authors actually conclude
Adam Rodman, one of the lead authors, told reporters there is “no formal framework right now for accountability” around AI diagnoses, and the study explicitly says the findings do not show AI is ready to make clinical decisions alone. The team is asking for real-world prospective trials before the technology is used at the bedside without a doctor in the loop. They also note the study only tested text inputs — it didn’t include images, audio, or any of the non-text information that real ER work depends on.
The framing the authors prefer is “AI as a colleague that catches what tired humans miss.” That’s a useful framing for a busy ER, where mistakes happen because people are stretched thin and pattern-recognition slows down at hour eleven of a twelve-hour shift.
The radiologist precedent
People predicted in 2017 that AI would replace radiologists by 2022. It didn’t. There are now more radiologists working in the United States than there were before the AI hype cycle started. What changed is that AI now does some of the work that used to take radiologists hours (initial flagging of likely abnormalities, first-pass measurement of tumors, comparison against prior scans) so the human expert can spend their time on the harder calls.
The same pattern is the most likely outcome here. Not “the AI replaces the ER doc,” but “the AI catches the lupus the doc would have missed at 3 AM, and the doc still makes the call.” That’s a better outcome for patients than either pure-human or pure-AI diagnosis.
What this means for patients
Three concrete shifts likely follow from this kind of evidence. First, fewer missed diagnoses for rare or atypical conditions. Lupus is the example everyone will quote, but the same principle applies to anything that requires a doctor to recall a textbook condition under time pressure. Second, faster triage decisions, because the AI’s first-pass differential takes seconds. Third, fewer unnecessary tests, because the AI can flag likely diagnoses early and reduce the “shotgun panel” approach.
What it does not mean is “the AI will see you instead of the doctor.” Hospitals are heavily regulated, malpractice law has not caught up to AI diagnoses, and the authors themselves are clear that the second-opinion role is what’s defensible right now.
What still doesn’t work
A few hard limits remain. AI models still hallucinate — produce fluent, confident text that’s flat-out wrong — and a confidently wrong diagnosis can do more damage than no diagnosis at all. They handle text well but struggle with the messy multimodal reality of an ER (a glance at the patient, a tone of voice, a smell, a vital-signs trend on a monitor). They don’t know when they don’t know, which is the single most important skill in emergency medicine. And the regulatory framework for “AI-generated diagnosis” is, at best, half-built.
Until those gaps close, “second opinion” is the responsible deployment. Which, conveniently, is also where the technology is genuinely useful right now.
The bottom line
An AI model just outperformed two doctors on a real-world diagnostic task. That’s a real result and worth taking seriously. It’s also not a replacement story — it’s a colleague story, and the authors are the ones saying so. Beginners in AI covers the AI-in-medicine beat in plain English, including which tools are FDA-cleared, which are research-only, and which are vendor hype — subscribe to the daily newsletter to keep up.
Sources and further reading
- TechCrunch: Harvard study finds AI offered more accurate ER diagnoses (May 3, 2026)
- Manrai, Rodman et al., study published in Science (May 2026)
- What is an LLM? — Beginners in AI glossary
- Start here: the Beginners in AI learning path
Get Smarter About AI Every Morning
Free daily newsletter — one story, one tool, one tip. Plain English, no jargon.
Free forever. Unsubscribe anytime.
Two ways to go further
The AI Prompt Library
1,000+ ready-to-use prompts for Claude, ChatGPT, and Gemini. Stop staring at a blank box.
Get it for $39 →2-Hour Live AI Crash Course
A private, beginner-friendly session across Claude, ChatGPT, Gemini, and the wider landscape.
Book for $125 →