This paper synthesizes publicly available evidence on the role of large language models (LLMs) in cybersecurity operations. Rather than proposing new benchmarks, the work audits existing research—including Mozilla advisories, peer-reviewed security studies, and regulatory standards—to assess which cybersecurity tasks are currently well-supported, weakly supported, or not supported for LLM deployment. The authors argue that current evidence supports LLM-assisted vulnerability analysis, penetration-testing workflows, and incident-response augmentation, but does not support claims of fully autonomous security operations without substantial governance controls. The paper contributes a machine-readable evidence matrix and task-level summary to enable independent verification of its claims.
Plajutin et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: