BackUrology

A multimodal vision–language model for generalizable annotation-free pathology localization

NatureResearch Authors: Hao Yang, Hong-Yu Zhou, Jiarun Liu, Weijian Huang, Cheng Li, Zhihuan Li, Yuanxu Gao, Qiegen Liu, Yong Liang, Qi Yang, Song Wu, Tao Tan, Hairong Zheng, Kang Zhang, Shanshan WangAIIM Authors: Junhyeok Hong, Madison SchanzApproved by President Reda RiffiPublication Date: 10/24/2025

Comprehensive Summary

This paper introduced AFLoc (Actionable Feature Localization), which is a vision-language model designed to localize clinically relevant abnormalities in medical images without requiring expert-level annotations. The authors evaluated AFLoc across multiple imaging modalities, such as histopathology, chest X-ray, and retinal fundus images, using annotation-free localization and zero-shot diagnostic tasks. In histopathology, AFLoc was set up based on the Quilt-1M dataset and tested on the SICAPv2 dataset (n = 2,100), where it outperformed both general-purpose and pathology-specific models in localization performance. Similar performances were observed across chest X-ray and retinal imaging benchmarks, where AFLoc achieved strong diagnostic accuracy without task-specific adjustments. The study also demonstrated that using more clinically meaningful text prompts improved localization performance, highlighting the importance of clinical language in guiding model outputs. Overall, AFLoc outperformed existing methods in both annotation-free localization and classification tasks and showed strong performance in several settings. Its localization accuracy exceeded human benchmarks, highlighting its potential to reduce reliance on manual annotations while remaining effective in complex clinical environments.

Outcomes and Implications

The findings in this paper suggest that annotation-free localization models like AFLoc could reduce reliance on manually labeled datasets while still producing accurate results. This is especially important in histopathology, where accurate localization is critical for tasks like cancer grading. Its strong performance across different imaging modalities and disease types shows the potential for more generalizable AI systems that can adapt to diverse clinical settings. Furthermore, given that there were improvements in outcomes with clinically precise prompts underscores the role of clinician input in shaping effective outputs. AFLoc ultimately showed promising technical performance, though real-world clinical evaluation will be necessary to determine how it can best assist clinician interpretation in practice.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.