Wispr Flow turns your voice into clean, ready-to-send writing — speak naturally, it strips the filler and fixes the punctuation. I've used it daily since February 2026 to build this newsletter. Read the full review →
Ninety patients had already run out of medicine.
Every one had been referred to the Undiagnosed Diseases Network, the NIH program that takes the cases nobody else can crack. Every one had spent a median of 7.6 years getting there. In the typical case, the symptoms started at seven months old.
Researchers at Vanderbilt took those 90 cases, all of them eventually solved, and handed each patient intake summary to two AI models.
ChatGPT-4o named the exact final diagnosis in 13.3% of them. Meta's Llama 3.1 8B got 10.0%. The clinical-review benchmark for those same cases was 5.6%.
AI more than doubled the doctors. That is the headline running everywhere this week. It is also the least interesting thing in the paper.
The contest was not a fair fight
The study is a research letter, which means it is short. The place its central comparison actually gets tested is the comment thread underneath it.
A physician reading it asked the obvious question: was the 5.6% produced the same way as the AI figures? Same inputs, same time pressure, same blinded scoring?
The authors answered in November 2025. No. The 5.6% reflects historical performance from routine clinical care, carrying none of the timing constraints and none of the blinded adjudication the AI outputs were held to.
That is not a scandal. The authors said so plainly the moment they were asked, which is exactly how this is supposed to work. But the clean line — AI 13.3% against doctors 5.6% — is comparing two things measured differently. And the clean line is the one that travels.
Two further limits they name themselves: the models read intake summaries rather than complete workups, and every one of the 90 cases had already been solved. Nobody has tested these models on the cases that are still open.
1,000+ Claude Prompts Top Professionals Actually Use at Work
Claude can be your analyst, editor, and strategist.
But most professionals are using it to fix grammar.
These 1,000+ Claude prompts take it from grammar tool to your most powerful AI work assistant.
Sign up for Superhuman AI and get:
1,000+ ready-to-use Claude prompts to get real work done in minutes — researched, tested, and used by professionals at Google, Microsoft, and NASA
Superhuman AI newsletter (4 min daily) so you keep learning new AI tools and skills to stay ahead in your career — the prompts are just the beginning
Why 13% still counts
Three cents and five seconds. That is what ChatGPT-4o cost per case. The self-hosted Llama model cost nothing at all and took two minutes.
Nobody is proposing a chatbot replace a clinical geneticist. The realistic question is much smaller: is a free first-pass list of hypotheses worth generating before an expert opens the file? On a case that has already consumed seven years, a differential that is helpful roughly one time in four, for the price of nothing, is difficult to argue against.
The tools actually working are not chatbots
In April 2026 the FDA cleared an ECG algorithm from Anumana, a company co-founded by Mayo Clinic, to flag cardiac amyloidosis from a standard 12-lead ECG. In a validation study of more than 15,000 adults it caught 78.9% of cases while correctly clearing 91.2% of non-cases.
Face2Gene has been in clinical use for years. Its model was trained on more than 17,000 photographs spanning 216 genetic syndromes. On images it had never seen, its single best guess was right about 65% of the time, but the correct syndrome sat in its top ten about 91% of the time.
Notice the shape of that second number. The tool is not built to answer a question. It is built to hand a clinician a short list worth checking.
The ceiling is data, and it is uneven
There are more than 7,000 rare diseases affecting an estimated 400 million people, and most individual conditions have only a few hundred documented cases, many sitting in records nobody ever digitized.
The data that does exist is skewed. A 2017 study found Face2Gene recognized Down syndrome in about 80% of Caucasian faces and 36.8% of African ones. Retraining on African photographs narrowed the gap, and newer versions do better. The mechanism does not go away: a model inherits the demographics of whoever supplied its training photos. Every claim about these tools carries a quiet second question. Accurate for whom?
What to do with this
A chatbot is a hypothesis generator, not an answer. Used well it produces a short list to bring to a doctor. Used badly it produces a confident condition name that anchors everyone in the room, including the clinician, to the wrong idea. Bring questions, not a conclusion.
And the lesson travels well past medicine. Next time you see a number showing AI beating professionals at something, do not ask whether the number is real. It usually is. Ask whether both sides were measured the same way. Here the researchers themselves said they were not, and you had to read the comments to find out.
The site version adds the copy-paste prompt, an interactive that lets you sort five real claims into what they actually are, and a proof table with every number and its primary source.
Free email without sacrificing your privacy
Gmail tracks you. Proton doesn’t. Get private email that puts your data — and your privacy — first.
— Jerry



