Original Investigation

Large Language Model Performance and Clinical Reasoning Tasks

JAMA Network Open 10.1001/jamanetworkopen.2026.4003

April 13, 2026 at 11:00 AM EDT

Read the full article

Can off-the-shelf large language models (LLMs) demonstrate reliable performance across the clinical workflow?In this cross-sectional study of 21 frontier LLMs tested on 29 standardized clinical vignettes, Grok 4 and other reasoning-optimized models achieved the highest scores, while Gemini 1.5 Flash performed lowest. Differential diagnosis consistently showed the weakest performance, while final diagnosis and management had stronger performances.These findings suggest that despite progress, current LLMs remain limited in early diagnostic reasoning and cannot yet be relied on for unsupervised patient-facing clinical decision-making.

Corresponding Authors: Arya S. Rao, BA, Harvard Medical School, 25 Shattuck St, Boston, MA 02115 (arya_rao@hms.harvard.edu); Marc D. Succi, MD, Mass General Brigham, 55 Fruit St, Boston, MA 02114 (msucci@mgh.harvard.edu).

Link to the article in your story

We encourage you to link out to this article in your story using the link below. It includes an access token that will give free access to the article for your readers up to one year after publication. (The link will be live after the article publishes and embargo is lifted.)

Please see the article for additional information, including full author list, author contributions and affiliations, conflict of interest and financial disclosures, and funding and support.

Need more information? Contact us.

Editor's Picks