BackPublic Health

Identifying Early Signals From Emerging Public Health Events Using Natural Language Processing.

Interdisciplinary Perspectives on Infectious DiseasesResearch Authors: Kelly S Peterson, Christian Dalton, Andrea Kalvesmaki, JoAnn Vuong, Colton Gordon, Senthil Nachimuthu, Mary Jo Pugh, Makoto M JonesAIIM Authors: Pearl Marks, Amanda ZhongApproved by President Reda RiffiPublication Date: 3/6/2026

Comprehensive Summary

This retrospective, multicenter study asked whether early, nonspecific signals of emerging infectious diseases, specifically public health authority communication, zoonotic exposures, and other pathogen exposures, can be automatically identified in clinical text, using a combination of rule-based natural language processing (NLP) and transformer-based models to perform text classification and information extraction. Researchers analyzed approximately n = 33,809,595 emergency department visits (clinical notes and associated EHR data) from the U.S. Department of Veterans Affairs across 113 medical centers between 2004 and 2024. Preprocessing included sentence segmentation and manual annotation of text spans, with additional augmentation using synthetically generated examples from large language models. The models tested included a rule-based medspaCy system for public health communication and SetFit transformer-based classifiers (Sentence-BERT) for zoonotic and pathogen exposures, compared against manual human annotation as the reference standard. The best-performing models achieved positive predictive values (PPV) ranging from 0.615 to 1.0 across tasks, with variable sensitivity and F1 scores depending on the exposure category. The analysis showed that among annotated samples, 20.1% of sentences reflected confirmed public health authority communication, while zoonotic exposure classification yielded 2.7% affirmed, 0.1% negated, and 97.2% no exposure in broader inference datasets. For other pathogen exposures, nearly half of early COVID-19 cases and over one-third of mpox cases included documented person-to-person exposure, while diseases such as leptospirosis showed higher proportions of water and environmental exposures. Secondary analyses included distributional comparisons across diseases and keyword frequency analyses (e.g., “tick” and “rabbit” for tularemia; “mosquito” for dengue), demonstrating alignment between extracted exposures and known epidemiology. Additional results showed that PPV was highest for food and environmental exposures, while sensitivity varied across categories, indicating trade-offs in detection performance. Limitations include reliance on clinician documentation (introducing reporting bias), single-annotator labeling without inter-rater reliability assessment, class imbalance with many “no exposure” instances, and a predominantly older, male veteran population, limiting generalizability. External validation was not performed, and subgroup fairness analyses were absent. Findings reflect the ability to extract epidemiologic signals from clinical text rather than direct improvements in patient outcomes and do not imply clinical efficacy.

Outcomes and Implications

This study suggests that hybrid NLP approaches, combining rule-based systems and transformer models, can scale biosurveillance by extracting meaningful early warning signals from unstructured EHR data. Clinically, these methods could support earlier identification of emerging infectious threats by flagging unusual exposure patterns or public health interactions before formal diagnoses are established. In practice, such systems could be integrated into hospital surveillance pipelines to prioritize chart review or alert infection control teams, although real-world deployment will require external validation, improved sensitivity, and careful consideration of workflow integration and bias.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.