BackPublic Health

Diagnostic performance of four AI tools in pharmacology MCQs: Accuracy, sensitivity, and specificity

PLOS OneResearch Authors: Ayah J. Al-Rahahleh, Mai Z. Rizik, Fahmi Y. Al-Ashwal, Rana K. Abu-FarhaAIIM Authors: Rithu Girish, Amanda ZhongApproved by President Reda RiffiPublication Date: 12/16/2025

Comprehensive Summary

Al-Rahahleh et al. compare the sensitivity and accuracy of Microsoft Copilot, ChatGPT-3.5, Google Gemini, and DeepSeek AI in answering multiple-choice questions (MCQs) regarding the subject of pharmacology. The main aspects of pharmacology that were focused on included drug action, side effects, pharmacokinetics, and drug-drug interactions. To evaluate the model’s capabilities, researchers developed 80 validated MCQs pertaining to the cardiovascular, respiratory, gastrointestinal, and endocrine systems. The questions were written using established pharmacology textbooks and clinical databases, reviewed by two pharmacology professors, and pilot tested by pharmacy practitioners to determine the difficulty level. Each AI tool was asked the same questions using a standardized prompt, and performance was measured based on accuracy, sensitivity, specificity, and consistency of answers over two testing rounds. Microsoft Copilot demonstrated the strongest overall performance, achieving an accuracy of 87.5%, a sensitivity of 94.6%, and a specificity of 70.8%. ChatGPT-3.5 and Google Gemini showed similar accuracy levels of 76.3% and 75.0%, respectively. DeepSeek AI had the lowest accuracy and the lowest specificity, however, it showed the highest reproducibility, providing identical answers in 97.5% of cases between the two testing rounds. Across all four tools, accuracy decreased significantly as question difficulty increased. Al-Rahahleh et al. highlight that although AI tools, such as Microsoft Cipilot, demonstrate strong performance on basic pharmacology questions, their reduced accuracy on more complex items limits their reliability. These findings emphasize the importance of verifying AI-generated answers with trusted pharmacology resources and caution against using these tools as a primary input for clinical decision making.

Outcomes and Implications

Pharmacology studies the implications and effects of drugs on living organisms and is crucial for the manufacturing and administration of medications. Pharmacology directly affects patient safety by ensuring appropriate drug selection, determining potential side effects, and avoiding harmful drug use. Al-Rahahleh et al.’s research is significant because, as AI diagnostic tools become more common in clinical settings, it is increasingly important to evaluate their accuracy and reliability. The study’s findings suggest that AI tools can be useful for reinforcing basic pharmacology knowledge but are not yet reliable enough for independent clinical decision-making, especially given the observed decline in performance on complex questions and variability in reproducibility across models. Overall, while AI tools show promise in pharmacology-related diagnostic abilities, their limitations in accuracy, reproducibility, and performance on complex questions require cautious use in clinical settings.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.