Speech Emotion Recognition in Mental Health: Systematic Review of Voice-Based Applications.
JMIR PublicationsResearch Authors: Eric Jordan, Raphaël Terrisse, Valeria Lucarini, Motasem Alrahabi, Marie-Odile Krebs, Julien Desclés, Christophe LemeyAIIM Authors: Alisa (Basil) Aleksandrova, Layna ParaboschiApproved by President Reda RiffiPublication Date: 9/30/2025Comprehensive Summary
In this article, Jordan et al. present a systematic review of literature pertaining to the application of Speech Emotion Recognition (SER) in identifying psychiatric disorders. 3648 studies were screened, 85 were retrieved and assessed, and 14 were included in the final review. These studies included analysis of speech with an emotion recognition component within a clinical context. Of the included studies, 3 addressed suicide risk and suicide ideation, 8 analyzed depression and mood disorders, and 3 studied psychotic disorders. Throughout each of these categories, studies demonstrated good discrimination between patients and control groups, with many models achieving accuracies of 70-80% and AUC of approximately 0.8. Newer machine learning (ML) methods often performed better than older ones. These results show that SER techniques are promising within a clinical context and may be used to support the diagnostic process, monitor and predict critical symptoms, or assist in measuring treatment efficacy.
Outcomes and Implications
SER has many advantages; it is noninvasive, has potential for automated analysis, and can provide real-time objective assessments. Integrating SER into psychiatric evaluation may improve diagnostic accuracy, enhance early detection of mental health issues, and improve patient care. However, SER has several limitations, some of which include the lack of interpretability of ML models, ethical considerations surrounding the use of patient data, lack of ability to generalize SER algorithms across diverse populations, and inaccuracy of vocal recordings in clinical contexts. The article suggests future studies to investigate a speech to emotion and emotion to pathology approach where speech is analyzed by a SER system and the resulting emotional states are analyzed for pathology. The article also suggests the use of a dimensional approach to pathology and the use of multimodal approaches in future studies. Interdisciplinary collaboration is needed to refine this technology and make it viable for clinical implementation.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.