BackNeurology

Bias and generalizability of brain age prediction models: A multi-cohort evaluation with anatomical and interpretability insights

Imaging NeuroscienceResearch Authors: Lautaro J. Aguzin Parrilli, Martin A. BelzunceAIIM Authors: Sedra Mourad, Sara ElanchezhianApproved by President Reda RiffiPublication Date: 3/12/2026

Comprehensive Summary

This study evaluates whether brain age prediction models can reliably detect abnormal brain aging across different populations and imaging protocols. Researchers tested four publicly available brain age models (ENIGMA, DeepBrainNet, Pyment, and BrainAgeNeXt) on 1,634 participants across four independent MRI datasets including Alzheimer's disease patients, long COVID patients, and healthy controls. Models based on 3D convolutional neural networks (Pyment and BrainAgeNeXt) achieved the best accuracy with error rates between 3.7 and 3.9 years, compared to 6.2 to 12.4 years for the older ENIGMA and DeepBrainNet models. However, all models showed systematic age-related biases where predictions tended to overestimate brain age in younger individuals and underestimate it in older adults. When the researchers used machine learning explainability techniques to identify which brain regions drove the models' predictions, lateral ventricles (fluid-filled spaces in the brain) were consistently the most important feature. Importantly, group-level comparisons showed that Alzheimer's disease and mild cognitive impairment patients had significantly elevated brain age gap values compared to healthy controls, but individual patients could not be reliably assessed based on their brain age prediction alone due to these systematic biases.

Outcomes and Implications

This research is important because brain age prediction has emerged as a promising biomarker for detecting abnormal aging, but it cannot be reliably used to assess individual patients in clinical practice. Current models show substantial age-related biases that limit their ability to identify early brain aging in individual patients, particularly older adults, where biases are most pronounced. The study demonstrates that while brain age prediction works well for comparing groups (such as detecting group-level differences between Alzheimer's disease and healthy controls), it should not yet be used to screen or diagnose individual patients. Significant barriers prevent immediate clinical implementation. The models were trained on selected research cohorts that may not represent the broader population, and they showed statistically significant but small biases related to ethnicity, suggesting that underrepresented groups in neuroimaging research may receive less accurate assessments. Additionally, the study identified recruitment bias in Alzheimer's disease datasets where "supernormal" control subjects were selected, potentially skewing predictions. Before brain age gap could be used clinically, future models would need to address age biases through techniques such as training on more age-balanced datasets, applying post-hoc statistical corrections, or retraining models specifically for different age ranges. Research validation in representative populations and comparison against clinical outcomes would be required, likely taking 5 to 10 years before individual-level clinical use becomes feasible.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.