Evaluating the impact of data biases on algorithmic fairness and clinical utility of machine learning models for prolonged opioid use prediction
JAMIA OpenResearch Authors: Behzad Naderalvojoud, Catherine Curtin, Steven M Asch, Keith Humphreys, Tina Hernandez-BoussardAIIM Authors: Rainier Dippong, Layna ParaboschiApproved by President Reda RiffiPublication Date: 10/1/2025Comprehensive Summary
The study being conducted by Naderalvojoud et al. aims to determine how machine learning models in healthcare can unintentionally produce biased outcomes that can place high-risk individuals at an even greater risk. Traditional fairness evaluations (such as AUROC) are largely insufficient because they do not capture impacts of clinical decision making. This was a retrospective study that utilized observational data from the Stanford Research Repository. Researchers selected adult patients who underwent surgery from 2008 to 2019 and received at least 1 opioid within 30 days before or after surgery. The study proposed a three-stage evaluation framework, using internal validation of the original training data, external validation of a new population, and revalidation on the new population. The framework evaluated clinical utility as well, using a new fairness concept called standardized net benefit (SNB) parity, which measures whether different patient subgroups gain equal real-world benefit. Research found that major disparities emerged during external validation, particularly for opioid-exposed patients, patients with depression, and those with severe comorbidities. These high-risk groups experienced lower predictive performance, systematic miscalibration, and reduced clinical utility, requiring much higher thresholds for the model to be beneficial. Overall, the research indicates that fairness in healthcare machine learning should be defined with equitable clinical utility, not just equal performance metrics.
Outcomes and Implications
Most machine learning (ML) evaluations in healthcare stop at performance metrics like AUROC or AUPRC. This study goes further by evaluating how models perform across clinically relevant patient subgroups (e.g. patients with depression or other severe comorbidities) and whether the models provide actual clinical benefit across those groups. This can help to bridge the gap between algorithm development and clinical application, which can be used to assess if a model is clinically useful and equitable across patient subgroups. The study shows that even models with acceptable overall accuracy can underestimate or overestimate risk for certain patients, which can be dangerous, especially in opoid prescription and postoperative care. Additionally, traditional fairness evaluations often focus on race/gender, which, while important, can eliminate clinically relevant groups (such as those with depression or diabetes) and risk-defined groups (including those who have previously been exposed to opioids and those who are “opioid-naive”). While there have been great strides in ML research, the work is still in the stages of pre-implementation and evaluation, and not a ready-to-deploy clinical tool. Additional development, calibration, and validation across clinical populations is needed, but it is already beginning to offer important metrics that clinicians can interpret in decision thresholds.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.