Comprehensive Pediatric Health Risk Stratification Using an AI-Driven Framework in Children Aged 2 to 8 Years: Design and Validation Study
JMIR Medical InformaticsResearch Authors: Zhihe Mao and Jundan ChenAIIM Authors: Pearl Marks, Amanda ZhongApproved by President Reda RiffiPublication Date: 1/26/2026Comprehensive Summary
This prospective, single-center study asked whether an artificial intelligence–driven, multimodal framework could accurately stratify pediatric health risks earlier and more effectively than traditional statistical and machine learning approaches, using a hybrid transformer-based NLP (BioBERT) model combined with ensemble learning to perform risk prediction and stratification. Researchers analyzed n > 40,000 pediatric participants aged 2–8 years using multimodal data, including structured and unstructured electronic health record (EHR) data, parental questionnaires, wearable sensor inputs, and online consultation records, collected over longitudinal time periods and split into training (70%), validation (15%), and independent test (15%) sets. The NLP component was built on BioBERT, fine-tuned on more than 50,000 pediatric medical journal articles and 100,000 deidentified clinical notes, and evaluated on 500 dual-annotated clinical notes. Comparator models included logistic regression, Cox proportional hazards models, support vector machines, decision trees, random forest, gradient boosting, and a conventional deep learning model. The best-performing proposed model achieved an AUC-ROC of 0.85 (95% CI 0.82–0.88), area under the precision-recall curve of 0.70 (95% CI 0.65–0.75), sensitivity of 0.78, specificity of 0.80, and F1-score of 0.75, significantly outperforming all baseline models (DeLong test, P < .05). The analysis was consistent between model predictions and expert review, with 78% (78/100) agreement in manually compared online consultation cases. Most misclassifications occurred in cases labeled “equivalent,” whereas clearly differentiated cases demonstrated higher predictive agreement. Subgroup analyses demonstrated consistent performance across demographic strata. AUC-ROC values were 0.84 for children younger than 2 years, 0.85 for those aged 2–5 years, and 0.86 for those older than 5 years, with similar stability across sex and socioeconomic subgroups. Secondary analyses included explainability modeling using SHAP values to quantify feature contributions, ensemble weighting strategies, confusion matrix evaluation, and statistical comparison testing. Limitations include single-center design, potential data imbalance across age groups (particularly ages 3–5), reliance on available EHR and digital data streams, and the absence of large-scale external validation. Although subgroup analyses were performed, advanced fairness mitigation techniques remain under development. Findings reflect predictive accuracy within a curated dataset rather than direct improvements in patient-centered outcomes and therefore do not establish clinical efficacy.
Outcomes and Implications
This study suggests that transformer-based NLP combined with ensemble learning can integrate heterogeneous pediatric data into clinically interpretable risk strata, offering earlier identification of risks such as obesity and developmental delay. In clinical practice, this framework could be embedded within pediatric EHR systems to provide real-time risk alerts, support shared decision-making with caregivers, and guide targeted referrals or preventive interventions. At a population level, risk distribution dashboards could inform public health resource allocation and proactive screening strategies. However, translation to bedside care requires multicenter external validation, longitudinal outcome studies, and rigorous bias mitigation to ensure equitable performance across pediatric populations.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.