BackDermatology

Hybrid vision transformer and graph neural network model with region-adaptive attention for enhanced skin cancer prediction

Scientific Reports (Nature Portfolio)Research Authors: Aswani Dogga, Sivasubramanian R. & Shanthi SAIIM Authors: Artiom Butuc, Josh BronteApproved by President Reda RiffiPublication Date: 2/7/2026

Comprehensive Summary

This study proposes a novel Hybrid Vision Transformer-Graph Neural Network (ViT-GNN) framework integrated with Region-Adaptive Attention (RAA) mechanism for robust skin cancer classification. The architecture combines VIT’s global contextual modeling with GNN-based spatial relational learning to capture complex lesion morphology and inter-region dependencies. The model was evaluated on three benchmark deroscopic datasets–ISIC 2020(33,126 images), HAM10000(10,015 images), and PH2(200 images)–using focal loss and class-balanced sampling to mitigate dataset imbalance. On ISIC 2020, the proposed model achieved 94.3% accuracy and AUC-ROC of 0.97, outperforming ResNet-50, EfficientNet-B3, and Swim Transformer. Cross-dataset robustness averaged 92.8% accuracy, with statistically significant improvements (p<0.05; Cohen’s d>0.8). Ablation studies confirmed that combining ViT+RAA+GNN consistently outperformed reduced configurations, demonstrating complementary gains from lesion-focused attention and graph-based relational modeling.

Outcomes and Implications

The introduction of Region-Adaptive Attention (RAA) enhances clinical relevance by explicitly prioritizing diagnostically significant lesion regions (e.g., asymmetrical borders, irregular pigmentation) while suppressing background artifacts such as hair and normal skin texture. Quantitative explainability analysis using Grad-CAM and SHAP showed improved alignment with dermatologist-annotated masks, achieving IoU up to 0.66 and Pointing Game accuracy of 88.7%, significantly surpassing CNN and transformer baselines. Despite its hybrid complexity, the model maintains competitive real-time feasibility with 19.4 ms inference time and 35.2M parameters. This architecture represents an advancement toward interpretable, clinically deployable AI systems in dermatology by integrating global attention, relational topology modeling, and adaptive lesion-focused weighting. Future directions include validation across underrepresented skin tones, integration of 3D dermoscopic imaging, and model compression for embedded healthcare deployment.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.