BackEmergency Medicine

Domain-Specific and Computer-Vision-Driven Versus General-Purpose AI Models in PA-CXR Analysis: a Comparative Study with Emergency-Medicine Specialists

Journal of Imaging Informatics in MedicineResearch Authors: AIIM Authors: Emma Edwards, Zaid ShehryarApproved by President Reda RiffiPublication Date: 1/5/2026

Comprehensive Summary

This prospective, single-center, cross-sectional diagnostic accuracy study compared the performance of three artificial intelligence (AI) models: GPT-5, Xray-GPT, and Qure.ai, with that of emergency medicine specialists (EMSs) in interpreting posteroanterior chest x-rays (PA-CXRs). Forty multiple-choice PA-CXR cases derived from real emergency department presentations between June and December 2024 were evaluated, representing six diagnostic categories and stratified by diagnostic complexity and clinical severity. Thirty EMSs completed the test once, while each AI model was tested daily over 30 consecutive days using standardized English-language prompts and identical clinical inputs. Qure.ai achieved the highest overall accuracy (median 37.0 out of 40; IQR 36.8-38.0), significantly outperforming EMSs and both GPT-based models (all p < 0.001). GPT-5 and Xray-GPT demonstrated approximately 50% accuracy and were significantly less accurate than EMSs across most categories. Qure.ai showed superior performance in both routine and challenging cases, as well as in life-threatening and non-life-threatening presentations, and achieved comparable accuracy to EMSs in interpreting normal CXRs, perforated peptic ulcers, and cavitary tuberculosis lesions.

Outcomes and Implications

Accurate and timely interpretation of PA-CXRs is critical in emergency medicine, where rapid diagnostic decisions are required for high-risk conditions. This study demonstrates that a domain-specific, computer-vision-based AI system trained on large-scale radiographic datasets achieved diagnostic accuracy comparable to or exceeding that of experienced emergency medicine specialists, while general-purpose GPT-based vision-language models demonstrated significantly lower performance. The authors conclude that specialized, image-optimized AI architectures outperform general-purpose multimodal language models in PA-CXR interpretation, emphasizing that domain-specific training is critical for achieving high diagnostic accuracy in emergency imaging. The authors note that further external validation in diverse clinical settings is required to assess generalizability and robustness.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.