BackEmergency Medicine

Artificial intelligence vs. emergency physicians: who diagnoses better?

Journal of the Brazilian Medical AssociationResearch Authors: Ali İhsan Kilci, Ramazan Azim Okyay, Erhan Kaya, Muhammed Semih Gedik, Hakan Hakkoymaz, Murat TepeAIIM Authors: Chloe Ng, Zaid ShehryarApproved by President Reda RiffiPublication Date: 12/5/2025

Comprehensive Summary

Kilci et al. conducted a cross-sectional study in Turkey to compare the diagnostic accuracy and initial test selection of an emergency medicine specialist against Large Language Model (LLM) algorithms using simulated clinical cases. From December 2024 to February 2025, an expert committee of emergency medicine professors developed case presentations including patient demographics, chief complaints, medical histories, vitals signs, and significant positive or negative examination findings. Imaging and laboratory findings were excluded to evaluate the diagnostic performance based solely on clinical presentation. The case represented a range of conditions across all five Emergency Severity Index (ESI) levels. An emergency medicine specialist with over five years of experience was compared to three unmodified OpenAI models: ChatGPT-4, ChatGPT-4o, and ChatGPT o3-mini. Responses were considered correct if they provided an exact match or a clinically appropriate approximation. The diagnostic accuracy rates for the specialist, ChatGPT-4, ChatGPT-4o, and ChatGPT o3-mini were 92.0%, 97.0%, 99.0%, and 99.0%, respectively. While the Cochran-Q test indicated a general difference between the four groups (p = 0.014), and the McNemar test initially yielded p = 0.039 for the top-performing LLMs, these results were not statistically significant after applying a Bonferroni-corrected significance threshold of alpha = 0.008. Similarly, no significant difference was found in the accuracy of initial diagnostic test selection (p = 0.208).

Outcomes and Implications

The findings suggest that while these models perform with high diagnostic accuracy consistent with recent literature, there is no definitive evidence that LLMs currently exceed the capabilities of human experts in a statistically significant manner. Furthermore, the study highlights that AI performance may vary in real-world settings characterized by time constraints, incomplete data, and emotional factors that influence decision-making. Unlike physicians, LLMs cannot currently provide ethical or legal justification for their decisions, which remains a primary concern for clinical accountability. Important limitations include the lack of real-world generalizability due to simulated cases, the use of a single human expert, and the lack of transparency regarding the architecture and training data of the LLMs.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.