The newest generation of artificial intelligence tools can now outperform experienced physicians in diagnosing complex medical cases, according to a team of researchers from Harvard, Stanford, the Massachusetts Institute of Technology and the University of Alberta.
“Our study suggests that large language models have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials,” the researchers conclude, suggesting the models should be tested as valuable second-opinion tools alongside doctors in clinical settings.
In a paper published today in the journal Science, the team tested the performance of OpenAI’s o1-preview model against five of the most difficult sets of benchmark cases, including the New England Journal of Medicine’s clinicopathological case conferences, which have been used to test medical minds for 100 years.
They also used data from recent real-world emergency department cases at a hospital in Boston, asking the AI models for a diagnosis based on data available at three time points — initial triage, examination by an emergency room doctor and admission to hospital. Both the AI model and the doctors were asked at each stage to provide a “differential diagnosis” — the five most likely issues affecting the patient.
The team found that the o1-preview model was able to identify the correct diagnosis in 78.3 per cent of the historic cases and suggested a helpful plan for further tests. In the emergency department cases, the model’s differential diagnosis included the exact or a very close diagnosis in 67.1 per cent of cases at triage, 72.4 per cent after physician examination and 81.6 per cent at admission to hospital or ICU. It performed significantly better than the two attending physicians at the first two stages and similarly to the physicians at the third stage.
“We thought earlier models were maybe going to change medicine,” says co-author and U of A neurology resident Liam McCoy, who joined the team during a dedicated residency research period and helped with statistical analysis, interpretation and framing of results.
“The moment I saw these results, I knew for sure this is going to change medicine. We’re already at the point where — if thoughtfully deployed — these models could be very, very useful in human-machine collaboration.”
A “thinking” machine
McCoy notes that the OpenAI o1-preview is a “big step up” from previous AI models such as ChatGPT because it uses a cyclical “thinking” process, essentially talking to itself, weighing options and checking its own logic before delivering an answer. The model had been tested favourably in other fields such as mathematics and software engineering, so the researchers wondered how helpful it would be in clinical settings.
“This model really blew humans out of the water on a number of different tasks,” McCoy says. “That doesn’t mean it’s ready to replace doctors, but it offers a lot to improve the quality of medicine, and that’s really exciting.”
We thought earlier models were maybe going to change medicine. The moment I saw these results, I knew for sure this is going to change medicine. We’re already at the point where — if thoughtfully deployed — these models could be very, very useful in human-machine collaboration.
McCoy says his group’s new paper is “in dialogue” with a seminal article published in Science in 1959, which first described the potential of computers to help in medicine.
“That was the foundational work about how we should be thinking about the process of medical diagnosis,” he says. “Now, 67 years later, we have finally got to the point where the models they dreamed of are able to perform at least a good portion of this reasoning.”
Although describing that performance as “superhuman” on these specific tasks, McCoy points out a phenomenon known as the “jagged frontier,” which means AI still also makes dumb mistakes — and sometimes even harmful recommendations.
Making medicine more human
“In certain tasks we see that the models have weaknesses and reasoning limitations,” he says. “The challenge of this next era is both to work on the technical level to expand the things they’re good at and shrink the things they’re bad at, and then also to work on the practical implementation level to understand, ‘OK, are we using models on the tasks they’re the best at, or are we using models on tasks where they fail?’”
Some people are justifiably afraid that using AI might be less humanistic, or alienating, or brought in just as a cost-cutting measure. But I really think there are numerous ways to get creative so we can make medicine more human and more caring with these tools.
McCoy will join the faculty at Beth Israel Deaconess Medical Center once he completes his residency at the University of Alberta Hospital in a year and a half. The team there has funding to begin testing what’s known as “collaborative teaming,” with doctors working alongside AI to see how their second opinions can improve patient outcomes.
Eventually, he expects to see clinical trials that evaluate the AI tools in real-life hospital settings.
“You would have some physicians who have access to this tool, some physicians who don’t, and you ask, ‘Are they making better diagnoses? Are the patients more satisfied? Are we reducing mortality? Are we reducing delays of appropriate treatment?”
He points to successes such as a newly approved AI tool now in use in the United Kingdom to diagnose stroke more quickly and the U of A’s own “Jenkins” AI scribe tool, which is now being tested across Alberta.
He sees potential from AI in nearly every medical field, from training students (creating study notes and finding research) to palliative care (imagine an AI chatbot that could answer repeated questions about your diagnosis whenever you want to ask them, even in the middle of the night) and communication disorders (AI could “listen” to patients with aphasia and fill in the gaps in their speech for others).
He foresees a time very soon when — once these tools have been validated in prospective trials and safely integrated into clinical workflows — it will be considered unethical not to use AI, much as other powerful technologies such as MRI have become part of standard care.
“Some people are justifiably afraid that using AI might be less humanistic, or alienating, or brought in just as a cost-cutting measure,” he says. “But I really think there are numerous ways to get creative so we can make medicine more human and more caring with these tools.”