Key points are not available for this paper at this time.
Medical artificial intelligence systems often rely on a single model, which may give inconsistent answers and does not reflect the team-based nature of clinical consultation. This study develops a collaborative framework in which several agents analyze a medical question from different professional perspectives, discuss the evidence, vote on the proposed answer, and revise unresolved questions up to three times. Each step and stopping condition is recorded in a time-ordered workflow so that the decision process can be reviewed. The framework was evaluated on three regional versions of a medical examination dataset and three medical image question-answering datasets. It achieved an accuracy of 80.1% on the Mainland examination dataset and produced the highest reported results among the evaluated systems on ADAM. Stronger image models performed better on ACRIMA and Covid CT, showing that collaboration remains limited by the underlying visual model. Additional model-based quality scores are reported only as exploratory results because they were not validated by an independent model or human experts. These findings show that structured collaboration can make medical artificial intelligence workflows more transparent.
Zhang et al. (Thu,) studied this question.