Salus: Strategic Diagnostic Testing for Complex Diagnosis via Multi-Agent Reinforcement Learning
Abstract
Lay Summary
Diagnosing a difficult illness is rarely a one-shot question. Doctors often start with incomplete clues, order tests, update their suspicions, and decide when they have enough evidence to act. Today's AI systems can answer many medical exam questions, but they often do poorly in this back-and-forth process: they may stop too early, choose the wrong tests, or miss rare but serious conditions. We build CompDiag-Bench, a testbed that turns complex diagnosis into an interactive task, and Salus, an AI system trained to reason through it. Salus splits the job into three cooperating roles: considering possible diseases, deciding whether more evidence is needed, and proposing useful tests. We then train these roles with feedback that rewards careful, evidence-based decisions rather than quick guesses. Our results show that a relatively small open model can diagnose complex cases more accurately than much larger systems in this setting, suggesting a path toward safer AI tools that assist, but do not replace, clinicians.