Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models
JAMA Ophthalmology 10.1001/jamaophthalmol.2025.2413July 31, 2025 at 11:00 AM EDT
What is the performance of OpenAI o1 compared with other large language models when addressing ophthalmology-specific questions?In this cross-sectional study among OpenAI o1 and 5 other large language models using almost 7000 ophthalmological questions from the Medical Multiple-Choice Question Answering (MedMCQA) dataset, o1 placed first in accuracy and macro F1 score but third in reasoning capabilities based on text-generation metrics; subgroup analyses showed that o1 performed better on queries with longer ground truth explanations. Human expert evaluation indicated that o1’s outputs were more clinically useful and organized than GPT-4o, but both provided high comprehensibility.The reasoning enhancements of o1 may not extend fully to ophthalmology, underscoring the need for domain-specific refinement in specialized medical fields.
Link to the article in your story
We encourage you to link out to this article in your story using the link below. It includes an access token that will give free access to the article for your readers up to one year after publication. (The link will be live after the article publishes and embargo is lifted.)
Please see the article for additional information, including full author list, author contributions and affiliations, conflict of interest and financial disclosures, and funding and support.
Need more information? Contact us.