Artificial intelligence versus human expertise: ECG-based detection of occlusive myocardial infarction after cardiac arrest
ResuscitationResearch Authors: Claudio Silwanis, Johannes Eder, Alexander Fellner, Alexander Nahler, Max Groche, Hermann Blessberger, Jörg Kellermair, Anna Neunteufel, Maximilian Huss, Julian Maier, Clemens Steinwender, Thomas LambertAIIM Authors: Soumya Halmandge, Zaid ShehryarApproved by President Reda RiffiPublication Date: 11/18/2025Comprehensive Summary
This paper evaluates how accurately artificial intelligence and human cardiologists can detect occlusive myocardial infarction (OMI) on post-cardiac arrest ECGs. THis is a setting where metabolic disturbances, global ischemia, and reperfusion injury commonly distort ECG signals. The study was a single-center study of 97 adults who were resuscitated from cardiac arrest between 2020 and 2023, and had both, a valid post-ROSC ECG and coronary angiography. Four diagnostic approaches were tested. The approaches included two senior cardiologists, a domain-specific deep neural network (Queen of Hearts, QoH), and two large language model chatbots (ChatGPT and EKG Analyst). Diagnostic performance metrics included AUROC, sensitivity, specificity, PPV, NPV, as well as F1 scores. QoH demonstrated the strongest discrimination for true acute coronary occlusion (TIMI 0), with an AUROC of 0.846 (95 % CI 0.752 - 0.939), sensitivity of 84 %, and a specificity of 66.7 %. Human experts performed moderately well with AUROC 0.735 (95 % CI 0.622 - 0.848) and sensitivity 92 %, though specificity was lower at 41.7 %. Large language model chatbots performed badly. ChatGPT had an AUROC of 0.456, and EKG analyst 0.474, each showing 100 % sensitivity but only 2.8 % and 1.4 % specificity respectively. This indicated a near total overdiagnosis. For the broader OMI definition (TIMI 0-2 or TIMI 3 + peak hs-troponin > 1000 ng/L), QoH again had the highest AUROC at 0.745, followed by human experts an AUROC of 0.635. ChatGPT and EKG analysts again showed extremely high sensitivity (98.36 % both) but extremely low specificity (2.8 % and 0 % respectively). The discussion emphasizes that QoH maintained great performance although lacking awareness of post-arrest clinical context, while human experts continued to outperform general purpose AI due to clinical reasoning and experience.
Outcomes and Implications
This paper highlights the major clinical need for being able to quickly identify which post-cardiac arrest patients have OMI, since timely coronary angiography can alter survival and neurological outcomes. The findings demonstrate that general large language models such as ChatGPT and EKG analyst are not suitable for clinical interpretation, especially in high stakes situations, as their near-universal “positive” classification provides no meaningful diagnostic discrimination and risks dangerous and misleading overdiagnosis. However, specialized AI tools like QoH, which are designed specifically for ECG-based OMI identification, show accuracy comparable to or even slightly surpassing the human experts, suggesting that they may support decision making when expert cardiology input is limited, such as in prehospital settings or smaller hospitals. While there isn’t a concrete timeline for clinical implementation, there is an emphasis that validated, domain-specific AI solutions may soon be incorporated into emergency cardiovascular care, whereas general purpose LLMs require substantial further evaluation before applying in clinical settings.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.