Evaluating deep learning sepsis prediction models in ICUs under distribution shift: a multi-centre retrospective cohort study
NPJ Digital MedicineResearch Authors: Fanny Tranchellini, Youssef Farag, Catherine Jutzeler, Lakmal MeegahapolaAIIM Authors: Ivan Chen, Thomas RenfrewApproved by President Reda RiffiPublication Date: 3/3/2026Comprehensive Summary
The study systematically evaluates deep learning models for early sepsis prediction in intensive care units (ICUs), with a focus on distribution shifting – data trained on data from a source hospital (vitals, patient demographics, lab values) being deployed elsewhere. Using a large multi-center retrospective cohort of over 216,000 ICU stays from three major databases (MIMIC-IV, eICU, and HiRID), researchers compared five deployment strategies — generalization, fine-tuning, retraining, target training, supervised domain adaptation, and fusion training—across multiple deep learning architectures, including convolutional neural networks, long short-term memory networks, and InceptionTime models. The analysis revealed that significant distributional differences exist between datasets, leading to reduced model generalizability when applied across institutions. Among deployment strategies, fine-tuning, despite its widespread use, consistently underperformed. Retraining and fusion training performed best in data-scarce and data-rich settings, respectively. Notably, supervised domain adaptation demonstrated the most stable and robust improvements in intermediate data regimes, enhancing both the “area under the receiver operating characteristic curve” and precision-recall performance. The study highlights that effective deployment of AI-based sepsis prediction models requires tailoring strategies to the availability of local data and understanding how domain shifts (data being used at external sites) can influence predictive models. Furthermore, the results suggest that the conventionally accepted fine-tuning deep learning model should be less relied upon.
Outcomes and Implications
The present study highlights the importance of addressing inter-facility differences when deploying artificial intelligence models for sepsis prediction in real-world ICU settings. By demonstrating that commonly used approaches like fine-tuning actually underperform compared to alternatives like retraining and supervised domain adaptation, the research provides practical guidance to clinicians, engineers, and healthcare systems on how to optimally implement predictive models across different sites and institutions. The findings emphasize that model performance is highly dependent on local data availability, suggesting that deployment strategies should be tailored, using retraining in data-scarce environments and domain adaptation when moderate amounts of local data are available, to maintain reliability and clinical utility. This has direct implications for patient safety, as poorly adapted models risk missed diagnoses or excessive false alarms, whereas properly adapted models can preserve early detection of sepsis and improve intervention timing. The study also informs researchers about the need to move beyond standard fine-tuning and to prioritize methods that explicitly account for cross-site variability. Barriers to implementation remain, including the retrospective nature of the analysis, dependence on high-quality structured ICU data, and limited validation in diverse and resource-constrained medical settings. The significance of this study is that it demonstrates strong potential for improving the generalizability of AI-driven sepsis prediction. Further work is needed to validate these strategies prospectively and develop an optimal strategy for adapting AI models to novel clinical sites.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.