Benchmarking Large Language Models Against Multidisciplinary Tumor Boards in Urological Oncology: Results from the Blinded, Prospective CONCORDIA Study
European Urology OncologyResearch Authors: Emily Rinderknecht, Maximilian Haas, Marco J. Schnabel, Anton P. Kravchuk, Christof Schäfer, Stephan Siepmann, Roman Mayr, Dominik von Winning, Jochen Grassinger, Christopher Goßler, Fabian Pohl, Peter J. Siska, Florian Zeman, Johannes Breyer, Anna Schmelzer, Christian Gilfrich, Sabine D. Brookman-May, Maximilian Burger, Matthias MayAIIM Authors: Junhyeok Hong, Madison SchanzApproved by President Reda RiffiPublication Date: 12/11/2025Comprehensive Summary
Rinderknecht et al. evaluated whether large language models (LLMs) could provide cancer treatment recommendations comparable to multidisciplinary tumor boards (MTBs) in urological oncology. The authors compared ChatGPT-4 and Claude 3.5 Sonnet with real MTBs using 110 structured case scenarios focusing on locally advanced or metastatic genitourinary cancers. All recommendations were assessed by uro-oncologists using the modified System Causability Scale (mSCS), to measure clarity and clinical reasoning. The average mSCS score for MTBs was 0.85. Claude 3.5 Sonnet scored 0.73, ending slightly below the predefined noninferiority threshold, while ChatGPT-4 scored lower at 0.66. Both models performed better in locally advanced cases than in metastatic ones, suggesting that case complexity influenced the AI performance. Failure analysis showed that most AI errors were due to deviations from guidelines or omission of key treatment options rather than misinterpretation of disease stage.
Outcomes and Implications
This study shows that publicly available LLMs are not yet reliable enough to match multidisciplinary tumor boards for complex oncological decision-making. While Claude 3.5 Sonnet came close in less complex cases, both models fell short in advanced disease scenarios where nuanced clinical judgment is required. However, these findings still suggest that AI may have a role as a supportive tool for early case review, education, or decision support in settings with limited access to MTBs, but not as a replacement. While MTBs are essential for cancer treatment planning, their use is often limited by the resources they require. With careful validation, system upgrades, and strict clinician oversight, these AI models have the capability to be safely integrated into high-stakes cancer care in the future.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.