BackPsychiatry

Identifying psychiatric manifestations in outpatients with depression and anxiety: a large language model-based approach

NatureResearch Authors: Shihao Xu, Yiming Yan, Yanli Ding, Feng Li, Shu Zhang, Haoyun Tang, Chao Luo, Yan Li, Hao Liu, Yu Mei, Wenjie Gu, Hong Qiu, Yong Wang, Jianyin Qiu, Tao Yang, Zike Wang, Qing Zhang, Haiyang Geng, Yunyun Han, Jun Shao, Nils Opel, Lidong Bing, Min Zhao, Yifeng Xu, Xun Jiang & Jianhua ChenAIIM Authors: Valerie Xian, Layna ParaboschiApproved by President Reda RiffiPublication Date: 12/5/2025

Comprehensive Summary

This study by Xu et al. investigates the potential of using large language models (LLMs) to identify psychiatric symptoms from audio recordings of psychiatrist-patient conversations and use them as intermediate features to predict diagnostic labels for depression and anxiety disorders in outpatients. Audio recordings from 1160 outpatients with depressive disorders or anxiety disorders were collected at the Shanghai Mental Health Center, yielding about 15,000 minutes of speech data. The researchers designed a corpus of 138 clinical indicators incorporating diagnostic criteria, main complaints, and mental status evaluations using Electronic Medical Records and assessment scales. They employed Qwen2-72B-Instruct as their foundational LLM, and trained it to identify clinical symptoms and rate components from six validated psychiatric rating scales (SCL-90, SDS, SAS, HAMD, HAMA, and MADRS). The LLM was also was fine-tuned with 477 high-quality psychiatrist annotations from EMRs The system achieved 86.9% accuracy for identifying the appearance of clinical annotations and above 70% accuracy for identifying symptoms of anxiety and depression. After supervised fine-tuning, the LLM's recall increased substantially from 66.1% to 81.1% on the test set and from 74.0% to 86.1% on the high-quality test set, while precision improved from 81.2% to 87.4%. For diagnostic classification, the model achieved a balanced accuracy of 75.5% with an AUPRC of 0.824 for distinguishing between anxiety and depression disorders, and 65.6% balanced accuracy for three-way classification (anxiety vs. depression vs. others). For symptom prediction, the model achieved 74.7% balanced accuracy for anxiety symptoms (sensitivity 0.683, specificity 0.810) and 77.2% balanced accuracy for depression symptoms (sensitivity 0.806, specificity 0.737). The study demonstrated that LLMs can effectively extract precise symptoms from psychiatric conversations for evidence-based diagnosis, achieving 86.1% recall after fine-tuning compared to 77.3% in zero-shot performance. Furthermore, the consistency of findings across different feature sets (clinical-related, assessment-related, LIWC, TF-IDF) strengthened reliability. This study addresses a critical gap by analyzing linguistic and symptom-related markers in clinical interview speech data from real-world, unstructured environments with first-episode outpatients, representing a significant advancement over controlled laboratory conditions and social media-based studies.

Outcomes and Implications

Depression and anxiety disorders affect over 300 million people globally each, with overlapping symptoms that make accurate diagnosis challenging. Current diagnostic approaches heavily rely on subjective observations constrained by time and clinical resources. While digital phenotyping offers promise for quantitative longitudinal observation, most existing studies rely on social media data or structured clinical reports with limited data availability, and there is a lack of research attempting to bridge the gap between patients' reported symptoms and professional diagnostic terminology used by clinicians. This research is demonstrates that LLM-generated features from clinical conversations can provide objective, data-driven insights for psychiatric diagnosis and assessment, potentially enhancing both efficiency and effectiveness. However, the authors note that real-world implementation demands careful consideration of practical factors including data privacy protocols, computationally capable infrastructure, clinician training for usability, ongoing validation and performance monitoring to detect model drift, and transparent ethical protocols including patient consent and equity audits. The authors emphasize that future work should include comprehensive symptom severity assessments, longitudinal analysis to track how linguistic patterns evolve with symptom progression or treatment response, validation with a larger and more diverse population, and expansion to a broader range of mental health conditions to ensure practical utility and generalizability.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.