The newest generation of artificial intelligence tools can now outperform experienced physicians in diagnosing complex medical cases, according to a team of researchers from Harvard, Stanford, the Massachusetts Institute of Technology and the University of Alberta.
“Our study suggests that large language models have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials,” the researchers conclude, suggesting the models should be tested as valuable second-opinion tools alongside doctors in clinical settings.
In a paper published today in the journal Science, the team tested the performance of OpenAI’s o1-preview model against five of the most difficult sets of benchmark cases, including the New England Journal of Medicine’s clinicopathological case conferences, which have been used to test medical minds for 100 years.
They also used data from recent real-world emergency department cases at a hospital in Boston, asking the AI models for a diagnosis based on data available at three time points — initial triage, examination by an emergency room doctor and admission to hospital. Both the AI model and the doctors were asked at each stage to provide a “differential diagnosis” — the five most likely issues affecting the patient.
The team found that the o1-preview model was able to identify the correct diagnosis in 78.3 per cent of the historic cases and suggested a helpful plan for further tests. In the emergency department cases, the model’s differential diagnosis included the exact or a very close diagnosis in 67.1 per cent of cases at triage, 72.4 per cent after physician examination and 81.6 per cent at admission to hospital or ICU. It performed significantly better than the two attending physicians at the first two stages and similarly to the physicians at the third stage.
“We thought earlier models were maybe going to change medicine,” says co-author and U of A neurology resident Liam McCoy, who joined the team during a dedicated residency research period and helped with statistical analysis, interpretation and framing of results.
“The moment I saw these results, I knew for sure this is going to change medicine. We’re already at the point where — if thoughtfully deployed — these models could be very, very useful in human-machine collaboration.”