Comparative feasibility of reasoning and non-reasoning large language models for gynecologic cancer emergency care
BMC Health Services ResearchResearch Authors: Junhwan Kim, Uisuk Kim, Jae Kyung Bae, Misong Kim, Ji Young Ham, Juwon Lim, Ji Hyun Kim, Sang-Soo Seo, Sokbom Kang, Sang-Yoon Park, Myong Cheol LimAIIM Authors: Ariyana Shafizadeh, Zaid ShehryarApproved by President Reda RiffiPublication Date: 11/24/2025Comprehensive Summary
Kim et al. sought out to determine the efficacy of large language models, more specifically the accessibility of the generative pretrained transformer (GPT)-4o and o3-mini-high and its role in assisting physicians diagnostically with patients diagnosed with gynecologic cancer. 15 cases were selected to evaluate these large language models, and their performances were compared to two gynecologic oncology fellows, two obstetrics and gynecology residents. They assessed the cases based on diagnostic measures utilized, conclusions drawn, diagnoses accuracy, and treatment options. Responses were assessed based on precision and speed. GPT-4o and o3-mini-high both responded more quickly than the physicians and were determined to be more accurate, with mean accuracy scores of 3.55 (95% CI, 2.98-4.10; P < 0.001) and 3.05 (95% CI, 1.98-3.88; P < 0.001), respectively. These findings suggest that large language models may have clinical utility in diagnostic evaluation when used cautiously and with appropriate oversight.
Outcomes and Implications
The findings suggest that large language models possess the capacity to assist clinicians in evaluating complex cases more quickly, generating thorough differential diagnoses, and outlining clear treatment plans. Such applications could reduce diagnostic delays, improve consistency, and ease the cognitive burden placed on providers. Despite their potential, these tools should complement, not replace, clinical judgment, and must be rigorously evaluated and validated before widespread implementation in clinical practice.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.