An agentic system for rare disease diagnosis with traceable reasoning
NatureResearch Authors: Weike Zhao, Chaoyi Wu, Yanjie Fan, Pengcheng Qiu, Xiaoman Zhang, Yuze Sun, Xiao Zhou, Shuju Zhang, Yu Peng, Yanfeng Wang, Xin Sun, Ya Zhang, Yongguo Yu, Kun Sun & Weidi XieAIIM Authors: Michaella Sevalie, Ahmad IslambouliApproved by President Reda RiffiPublication Date: 2/18/2026Comprehensive Summary
DeepRare examines whether an agentic large language model system can improve rare disease diagnosis using mixed clinical inputs. This was a retrospective, multicenter diagnostic benchmarking study drawing on nine datasets with 6,401 cases spanning Asia, North America, and Europe and covering 2,919 diseases. The system uses a multi-agent LLM architecture that combines free-text clinical notes, Human Phenotype Ontology terms, and genomic data. It was compared with traditional rare disease tools, standalone LLMs, and other agentic systems. On HPO-based tasks, DeepRare reached Recall@1 of 57.18%, compared with 33.39% for the next best method. When genetic data were added in the Xinhua cohort, Recall@1 increased to 69.1% versus 55.9% for Exomiser. Physician reviewers judged the model’s cited reasoning accurate in 95.4% of sampled cases. Some caution is warranted. Most testing was retrospective, and performance varied across specialties and datasets. In a review of 200 missed cases, the most frequent problems were phenotype weighting errors (82 of 200) and confusion between clinically similar disorders (77 of 200). Although the study included cross-center evaluation, detailed demographic fairness analyses were not reported, and confidence intervals were inconsistently provided. In practical terms, the system may help clinicians narrow rare disease differentials earlier in the workup, especially when both phenotype and genetic data are available. Prospective clinical studies are still needed before routine use.
Outcomes and Implications
If DeepRare performs the same way in real clinical settings, it could change how clinicians approach rare disease workups. Right now, many cases start with a very wide differential. A tool like this might help narrow that list earlier. That could be especially useful when symptoms span multiple organ systems or when genetic results are already available. There is also a workflow piece to consider. Because the system lays out its reasoning, clinicians can see how it arrived at a suggestion instead of just accepting the output at face value. That kind of transparency could be helpful in teaching environments or during team case discussions. Even so, we do not yet know how often clinicians would actually rely on these rankings in everyday practice. The failure patterns deserve some attention too. The model has the most difficulty when different conditions look very similar or when it needs to judge which clinical features carry the most weight. Those situations are challenging for human clinicians as well, which makes the limitation important rather than surprising. The study also does not provide detailed fairness analyses, so it remains unclear how consistent performance is across patient groups. At this stage, DeepRare fits best as decision support, not as a replacement for clinical judgment. What really matters next is prospective testing inside real clinical workflows. That will show whether the system meaningfully shortens time to diagnosis, changes testing behavior, or improves patient outcomes.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.