BackUrology

Artificial intelligence versus classical scoring systems: a comparative analysis of stone-free prediction after percutaneous nephrolithotomy

UrolithiasisResearch Authors: Burak Elmaağaç, Ali Yasin Özercan, Abdullah Gölbaşı, Hüseyin Biçer, Ercan Arslan, Mert Ali KaradağAIIM Authors: Junhyeok Hong, Madison SchanzApproved by President Reda RiffiPublication Date: 1/30/2026

Comprehensive Summary

Elmaağaç et al. compared a urinary stone–free rate predictor based on ChatGPT against an established nephrolithometric scoring system for predicting outcomes after percutaneous nephrolithotomy (PNL). The authors analyzed 340 patients treated between 2019 and 2025 and assessed stone complexity using Guy’s Stone Score (GSS), S.T.O.N.E., S-ReSC, and the CROES nomogram, alongside predictions generated by a ChatGPT-based tool. The overall stone-free rate was 60.9%. Patients who became stone-free tended to have less complex stones, which was reflected by lower GSS, S.T.O.N.E., and S-ReSC scores. These scoring systems were able to distinguish between patients who cleared their stones and those who did not, even after accounting for other factors. In contrast, predictions generated by the ChatGPT-based tool and the CROES nomogram did not meaningfully differ between the two groups, indicating limited usefulness. Overall, the ChatGPT-based model did not reliably predict surgical outcomes and performed little better than chance.

Outcomes and Implications

The results from this study show that traditional scoring systems remain more reliable than general-purpose AI tools for predicting surgical outcomes after PNL, particularly in complex stone disease. While ChatGPT is capable of integrating multiple clinical inputs, it lacks transparent training and disease-specific optimization. This limits its usefulness as a decision-support tool in endourology. The study highlights an important distinction between AI models designed for medical information support and those required for accurate surgical outcome prediction. Future progress in this area will likely depend on condition-specific machine learning models trained on structured clinical datasets, along with external and prospective validation, rather than reliance on general large language models.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.