Evaluating large language model performance in answering “Principles of Health” course questions
Scientific ReportsResearch Authors: Mohsen Khosravi, Emine Kübra Dindar, Burak SayarAIIM Authors: Fatema Dinary, Amanda ZhongApproved by President Reda RiffiPublication Date: 4/26/2026Comprehensive Summary
This research presented by Khosravi et al. conducted an empirical evaluation study of the efficacy and accuracy of four large language models in answering health science education questions. Exam-style questions were posed to ChatGPT-4o, Gemini 2.5, Microsoft Copilot, and Perplexity AI evaluating correctness by comparing against answer keys and expert scoring. The results found that ChatGPT-4o and Perplexity 2.250619.0 had the highest overall accuracy with 93% scores for each respectively while Gemini 2.5 and Copilot 2025 were each 86% accurate. In terms of specificity, ChatGPT-4o and Perplexity scored the highest (0.80) while Gemini 2.5 and Copilot 2025 had lowest results with 0.66 and 0.60 scores respectively. Khosravi et al. acknowledged the limitation of decreased performance as questions became more complex and the need for further research.
Outcomes and Implications
Given that large language models are being increasingly used by students to answer health science education questions, understanding the merits and shortcomings of using LLMs as academic support is becoming especially relevant. Incorrect AI responses, however, could detrimentally affect medical learning. As artificial intelligence is becoming more prominent in the STEM and medical field, this study helps to determine the reliability of generative language models.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.