Multimodal Large Language Models for Inflammatory Skin Disease Evaluation: A Cross-Sectional Study
Journal of the American Academy of DermatologyResearch Authors: Arjun Mahajan, Maureen Whittelsey, Keyarah Grullon, Evan W. Piette, Jeffrey A. Sparks, David W. Bates, Vinod E. Nambudiri, Avery H. LaChanceAIIM Authors: Hanna Zhu, Josh BronteApproved by President Reda RiffiPublication Date: 2/5/2026Comprehensive Summary
This study investigates how well multimodal large language models (mLLMs) can diagnose inflammatory skin diseases using clinical images. Researchers evaluated GPT-5, Gemini-2.5-Pro, and Janus-pro-7b on 1758 dermatologist-labeled images from the Stanford and SCIN datasets, with 12 different inflammatory skin conditions and diverse Fitzpatrick skin types. They were evaluated using standardized prompts and diagnostic accuracy with confidence intervals was calculated. Overall, results show that GPT-5 and Gemini performed better than Janus. Accuracy varied between different conditions, with GPT-5 performing the best for granuloma annulare but poorly for erythema multiforme. All models showed lower performance in darker skin tones (FST 5–6), younger age groups, and certain anatomic regions such as the foot. The authors state that while mLLMs are useful, they show biases related to skin tone, age, and lesion characteristics. The current systems are not able to be fully relied on for inflammatory disease classification.
Outcomes and Implications
This study highlights both the promise and limitations of AI in dermatology. This research is important because inflammatory skin diseases account for a large proportion of dermatology visits, but AI validation in this area is limited. mLLMs have the potential to be used as screening tools, but clinicians need to supervise the diagnosis to prevent error.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.