Evaluating large language models for specialist referral triage in primary care: a quantitative study using otolaryngology scenarios
Family PracticeResearch Authors: Hack S, Attal R, Yogev D, Alsleibi S, Gvili B, Shahal D, Glikson EAIIM Authors: Katharina Staehr, Zaid ShehryarApproved by President Reda RiffiPublication Date: 11/28/2025Comprehensive Summary
Hack et al. assessed whether two artificial intelligence (AI) models could effectively and reliably assist primary care providers in deciding when and where to refer patients to ear, nose, and throat (ENT) specialists. Researchers presented 16 simulated clinical scenarios (such as sudden hearing loss or neck masses) to both ChatGPT and Gemini in two formats: structured clinical prompts written by doctors, and informal queries written in patient language. Outputs were rated by five ENT specialists and ten lay reviewers based on appropriateness, clarity, safety, and usefulness. The performance of each model was compared as well as the agreement among reviewers and impact of prompt structure on output quality was assessed. The study found that when given structured clinical prompts, both models produced safe and clinically appropriate recommendations, with no statistically significant differences between models across any assessment metric. Notably, performance of both models decreased when given informal prompts. Reviewer agreement was strong (Intraclass Correlation Coefficient (ICC) 0.66–0.79) suggesting high reliability.
Outcomes and Implications
This study demonstrates that AI models with clear, structured prompting show meaningful potential in supporting specialist referral triage in primary care. Such models may help reduce care delays, unnecessary referrals, and strain on clinical systems. However, the study is limited by its reliance on simulated text-based scenarios, a small reviewer sample, restricted input formats, and the use of time-specific model versions, which may affect reproducibility. As such, physician oversight remains crucial, particularly in high-stakes clinical cases. Future prospective studies across diverse clinical settings are warranted to establish broader generalizability.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.