Are clinical improvements in large language models a reality? Longitudinal comparisons of ChatGPT models and DeepSeek-R1 for psychiatric assessments and interventions
Sage JournalsResearch Authors: Alexander Smith, Michael Liebrenz, and Ana BuadzeAIIM Authors: Valerie Xian, Layna ParaboschiApproved by President Reda RiffiPublication Date: 7/31/2025Comprehensive Summary
This study examines the diagnostic accuracy, treatment proposals, patient interactions, and cultural sensitivity of newer large language models including ChatGPT-4o, ChatGPT-4.5, and DeepSeek-R1 for psychiatric assessments and interventions. In the study, three psychiatric cases from earlier literature about sleep-related problems and co-occurring issues were used, allowing for cross-comparisons with a 2023 ChatGPT model baseline. The researchers evaluated all models using their March 2025 versions and assessed multiple dimensions including primary diagnosis accuracy, clinical reasoning, pharmacological recommendations, non-pharmacological advice, empathy in communication, risk evaluation capabilities, and cultural adaptations across the presented cases. They found that ChatGPT-4o, ChatGPT-4.5, and DeepSeek-R1 showed modest improvements from the 2023 ChatGPT model but still exhibited significant limitations. Communication was empathetic and non-pharmacological advice typically adhered to evidence-based practices. Primary diagnoses were generally accurate but often omitted somatic factors and comorbidities. Clinical reasoning worsened as case complexity increased, which was especially apparent for suicidality safeguards and risk stratification. Pharmacological recommendations frequently diverged from established guidelines, and cultural adaptations remained largely superficial. Output variance was noted in several cases, and the LLMs occasionally failed to clarify their inability to prescribe medication. Notably, DeepSeek-R1 performed comparably to ChatGPT models despite being open-source and developed with different architectural principles. Despite incremental advancements, ChatGPT-4o, ChatGPT-4.5 and DeepSeek-R1 were affected by major shortcomings, particularly in risk evaluation, evidence-based practice adherence, and cultural awareness. The tools cannot substitute mental health professionals but may still have benefits.
Outcomes and Implications
Potential clinical applications for emerging large-language models are well-documented, and newer systems like DeepSeek have attracted increasing attention, but important questions remain regarding their reliability and cultural responsiveness in psychiatric settings. Given the rapid evolution of AI technology and an increasing interest in using it for mental health support, longitudinal evaluation is essential to track whether claimed improvements in successive model versions translate to meaningful clinical gains, particularly for complex psychiatric assessments involving suicide risk and culturally diverse populations. This research is highly clinically relevant as it provides evidence that LLMs at the time of writing are insufficient for independent psychiatric practice due to persistent deficiencies in risk stratification, guideline adherence, and cultural sensitivity. The findings support using these tools only in adjunctive roles under professional supervision rather than as standalone diagnostic or treatment systems. The study suggests ongoing monitoring of successive model versions will be necessary to determine if or when these tools reach appropriate clinical readiness. Furthermore, the authors emphasize that substantial improvements in transparency, prompt engineering, risk evaluation capabilities, and cultural competence are prerequisites before these technologies can be safely and equitably deployed in psychiatric practice.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.