Recent advances in artificial intelligence have significantly improved performance in medical imaging tasks such as segmentation, quantification, and report generation. However, most existing solutions operate as static pipelines with limited adaptability, quality assurance, and workflow-level reasoning. In this work, we present ARGUS, an agentic framework for multimodal medical image analysis that coordinates specialized agents within a unified architecture. An Orchestrator Agent interprets user requests, identifies the imaging modality, and assembles task-specific execution plans by selectively engaging processing, quantification, verification, knowledge retrieval, and reporting agents. This enables context-aware decision-making and dynamic workflow reconfiguration based on intermediate findings and runtime conditions. A key feature of ARGUS is its ability to supervise and contextualize analytical processes. The Verification Agent performs quality control by assessing intermediate artifacts against task-specific criteria, while the Knowledge Retrieval Agent enriches quantitative findings with evidence from the biomedical literature and established physiological reference ranges. Together, these components promote transparency, support automated error detection, and reduce the risk of propagating unreliable information through downstream stages. The framework was evaluated across three imaging domains: radiology (MRI), pathology (hematopathology), and ophthalmology (OCT). Quantitative evaluation demonstrated strong agreement between ARGUS and reference standards across pathology, OCT, and MRI tasks, achieving a cell-counting bias of −0.182 cells (MAE = 0.727), a full retinal thickness bias of −31.30μm (MAE = 37.63μm), and MRI volumetric errors below 3mL, while also achieving closer agreement with reference measurements than the evaluated general-purpose and domain-specific baseline systems. These results demonstrate the feasibility and potential value of agent-based orchestration for enabling adaptive, validated, and interpretable multimodal imaging workflows while providing a scalable foundation for complex multi-step clinical analysis.
Helmy et al. (Tue,) studied this question.