AI in Patient Care: Evaluating Large Language Model Performance Against Evidence-Based Guidelines for Pulmonary Embolism
Thoracic Research and PracticeResearch Authors: Ömer F Karakoyun, Halil E Koyuncuoğlu, Ömer H Sağnıç, Mehmed E Özdemir, Yalçın Gölcük, Birdal YıldırımAIIM Authors: Husayn Ladha, Amine NoureddineApproved by President Reda RiffiPublication Date: 1/20/2026Comprehensive Summary
Karakoyun et al. conducted a controlled, case-based evaluation of four widely used large language models (ChatGPT-4o, Gemini, Grok, and DeepSeek-V2) to assess their ability to apply the 2019 European Society of Cardiology guidelines for the diagnosis and management of acute pulmonary embolism (PE). A single simulated high-risk PE scenario was used to generate ten open-ended questions spanning diagnosis and initial evaluation, risk stratification and prognosis, management and treatment, and post-treatment assessment and follow-up. Responses were generated under standardized prompting conditions and independently scored by two emergency medicine physicians using predefined, guideline-based reference answers on a 10-point scale per question (maximum score 100). ChatGPT-4o achieved the highest score (76), followed by Gemini (73.75), Grok (71.25), and DeepSeek-V2 (65), with no statistically significant difference in overall performance between models (Kruskal-Wallis H = 3.013, P = 0.390). Domain-specific performance varied, with DeepSeek-V2 performing relatively better in diagnostic components, Gemini in treatment planning, ChatGPT-4o in post-treatment assessment and follow-up, and Grok providing more pragmatic but less systematically guidelines-referenced responses. Inter-rater reliability was excellent (ICC: 0.986, 95% CI: 0.975-0.992), supporting the consistency of expert scoring.
Outcomes and Implications
This study demonstrates that current LLMs can generate broadly guideline-aligned responses for acute PE but exhibit meaningful variability across clinical domains and models. No system consistently performed well across diagnosis, risk stratification, treatment, and follow-up: expert reviewers identified recurring limitations, including incomplete risk stratification, insufficient patient-specific reasoning, and gaps in therapeutic detail. Given the single-scenario, simulation based design and the rapidly evolving nature of LLMs, these findings should be interpreted as proof-of-concept rather than evidence of readiness for autonomous clinical use. At present, LLMs may serve as complementary point-of-care reference tools to support clinician reasoning, but they should not replace physician judgement without prospective validation.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.