Automated Echocardiographic Detection of Congenital Heart Disease Using Artificial Intelligence
CirculationResearch Authors: Platon Lukyanenko, Sunil Ghelani, Yuting Yang, Bohan Jiang, Timothy Miller, David Harrild, Nao Sasaki, Francesca Sperotto, Danielle Sganga, John Triedman, Andrew J. Powell, Tal Geva, William G. La Cava, Joshua MayourianAIIM Authors: Husayn Ladha, Amine NoureddineApproved by President Reda RiffiPublication Date: 3/28/2026Comprehensive Summary
Lukyanenko et al. developed and tested EchoFocus-CHD, a view-agnostic multitask AI model that analyzes complete transthoracic echocardiograms for 12 critical and 8 non-critical congenital heart lesions. The study used a large pediatric echocardiography dataset, including 54,727 internal studies and 3,356 referral studies, totaling about 3.6 million videos overall. The referral cohort differed significantly from the internal cohort, with a higher prevalence of critical CHD (29.4% vs 5.8%) and younger patients overall. For the primary endpoint of composite critical CHD, internal performance was strong, with an AUROC of 0.94, a positive likelihood ratio (LR+) of 7.50, and a negative likelihood ratio (LR−) of 0.14. Performance worsened in outside datasets, with AUROC falling to 0.77 in the referral cohort. Calibration also worsened, with the scaled Brier score declining from 0.405 internally to 0.045 in the referral cohort. After retraining on a broader US dataset, international referral performance improved to an AUROC of 0.87 overall and 0.84 in infants, suggesting that greater training diversity partly reduced the model’s generalizability problem.
Outcomes and Implications
EchoFocus-CHD is clinically applicable as a triage tool rather than a stand-alone diagnostic system. Its internal negative likelihood ratio of 0.14 suggests good rule-out performance for critical CHD in the internal test cohort, which could help prioritize higher-risk studies for faster pediatric cardiology review. This may be particularly useful in tele-echocardiography networks, community hospitals, and resource-limited settings where expert interpretation is not immediately available. Nonetheless, the drop in AUROC from 0.94 internally to 0.77 in referral studies, together with the notable decline in calibration, indicates that the model is not yet appropriate for routine deployment across centers with different imaging practices and case mix. Clinically, this supports use as a workflow support tool rather than a replacement for expert interpretation.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.