Human oversight in clinical artificial intelligence is commonly concentrated at output release. In clinical variant classification, however, an autonomous component can influence the decision earlier by shaping the reference-data and evidence snapshot consumed by the classifier. This Perspective argues for a deterministic, model-output-free classification core that produces a frozen ACMG/AMP class and criterion trace, with language-model agents confined to a downstream interpretation layer. Within a five-gate framework, evidence admissibility (G1) is the load-bearing human transition: an authorized person promotes a provenanced, version-locked evidence snapshot before it can enter the live classifier, while qualified sign-out (G4) remains necessary at release. The proposed framework is operationalized through prospective metrics: unauthorized evidence-base admission rate, longitudinal change-attribution rate, model-to-core boundary violations, exact reconstruction rate, narrative contradiction rate, and prioritization omission rate. Under matched classifier, case-input, observation-window, and sign-out conditions, human-authorized evidence admission should reduce unauthorized change and unattributed drift and improve reconstruction of why a classification changed. The framework addresses auditability and governance. It does not establish correctness, clinical validation, comparative effectiveness, or fitness for purpose, and complete runtime separation increases the workload of human literature curation. Version note: This record is a substantially revised Version 2.0 of the preprint originally published at https://doi.org/10.5281/zenodo.21189571. The revision preserves the original architecture and central thesis, reorganizes the manuscript as a focused Perspective, adds prospective evaluation metrics, consolidates the figures, and calibrates the scientific claims. No new dataset or completed validation results are reported.
Vladimir Mitev (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: