BackMedical Informatics

Prospective quantitative analysis of hyperparameter and input optimization in GPT-5: comparative contribution to radiologist performance in abdominal radiology

Diagnostic and Interventional RadiologyResearch Authors: Eren Camur, Turay Cesur, Yasin Celal Gunes, Muhammed Batuhan Gokhan, Riza Sarper OktenAIIM Authors: Jiya Dave, Ahmad IslambouliApproved by President Reda RiffiPublication Date: 2/16/2026

Comprehensive Summary

This cross‑sectional study investigated how large multimodal AI models, particularly GPT‑5, can assist abdominal radiologists. Researchers evaluated how adjusting two hyperparameters, temperature (response creativity) and top‑p (response precision), affects diagnostic accuracy. Abdominal imaging cases from the Eurorad database included visual images, imaging findings, and clinical information, allowing comprehensive decision‑making. Browser‑based GPT‑5 and OpenAI’s interface were prompted to act as abdominal radiologists, providing one primary and four differential diagnoses. Varying temperature values and top‑p values were tested to find the optimal combination. The performance of two radiologists with and without GPT-5 was compared with that of an abdominal radiologist with 25 years of experience. Accuracy was measured by a binary correct/incorrect score for final diagnoses and a five‑point Likert scale for differentials. Results showed that as temperature and top‑p increased, GPT‑5’s diagnostic accuracy improved. The optimal combination was temperature = 1.5 and top‑p = 1, with a 73 % accuracy and an average 3.84 on the Likert scale. When clinical information was added to visual input, AI accuracy rose from 12 % to 58 %. Radiologists alone performed at 71–73 % accuracy; with GPT‑5 assistance, this increased to 86–87 %. Using GPT-5’s optimal settings, overall diagnostic accuracy for the two radiologists reached 94 % with a 4.56 and 4.49 score, slightly exceeding the experienced abdominal radiologist’s 92 % with a 4.4 score. Limitations include use of non‑clinical data, limited generalizability, single‑prompt testing, and possible variability or learning bias within the ChatGPT model over time. Overall, optimizing GPT‑5 settings significantly enhanced diagnostic precision and consistency, and can be especially useful for complex or atypical cases.

Outcomes and Implications

This research highlights the potential of multimodal AI systems, combining text and image training across extensive datasets, to strengthen radiologic decision-making. Because radiology depends on precision, optimizing AI models through key hyperparameters can improve both accuracy and contribute to better workflow efficiency. By comparing new patient cases with large, integrated databases of similar findings, these systems can support faster and more consistent diagnostic reasoning. Clinically, this improvement carries meaningful implications as a structured and accurate differential diagnosis list can help radiologists and physicians focus on the most likely conditions, reduce confidence uncertainty, and guide more effective patient management.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.