Background Agentic artificial intelligence systems capable of autonomous reasoning, planning, tool invocation, and clinical action are entering health services faster than existing governance instruments can accommodate. Whether empirical evaluations document the institutional arrangements these systems require has not been systematically examined across clinical domains. Methods Following PRISMA 2020, we searched OpenAlex, PubMed/MEDLINE, Scopus, and IEEE Xplore on 8 September 2026 for empirical studies of agentic artificial intelligence in clinical tasks, benchmarks, and deployment settings. Eligibility under the SPIDER framework required a clinical sample and evidence in the full text of at least three of four agentic capabilities: multi-step reasoning, planning, runtime tool-calling, and autonomous execution. Accountability findings reflect documentation within empirical evaluations rather than legal or regulatory literature. Two reviewers independently screened 1,793 records and assessed 81 full texts, with all eligibility disagreements formally adjudicated against source texts. Appraisal used a custom eight-domain codebook. Extraction covered 24 fields across four research questions addressing architecture, clinical safety, safeguards, and accountability. Results Fifty-two studies met inclusion criteria; 34 peer reviewed, 18 preprints, 32 from 2026. Across the predefined reporting domains, documentation was highest for architecture (96 per cent) and progressively less frequent for clinical safety (78 per cent), safeguards (62 per cent), and accountability (28 per cent). Vendor accountability and patient contestability each appeared in one study, workforce impact in three. Adversarial security was evaluated in five studies. Twenty-seven of 52 studies did not report a fail-safe protocol. Conclusions The evidence base describes agentic clinical systems in architectural detail and increasingly in safety terms, but provides limited empirical documentation of procurement standards, liability allocation, workforce assessment, and patient recourse. Safety evaluation and accountability instrumentation answer different questions, and closing the second is not a by-product of improving the first.
No takes yet. Share an insight, caveat, or question.
Muqoddam et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: