Large language model (LLM) watermarking is a prominent approach for identifying machine-generated text, supporting provenance, protecting intellectual property, and enabling accountability in generative-AI ecosystems. Proposed mechanisms include token- and logit-based sampling, semantic and syntactic watermarking, post-hoc methods, publicly verifiable schemes, model- and service-level watermarking, and retrieval-augmented or application-specific approaches. However, their practical reliability under realistic transformations, adaptive attacks, limited model access, diverse deployment environments, and high-stakes governance settings remains uncertain. This article presents an evidence-centered, PRISMA-guided systematic literature review based on searches completed on 14 May 2026. Rather than only classifying mechanisms, we assess how convincingly existing studies support their reliability and robustness claims. We identified 1367 records, retained 1065 after duplicate removal, assessed 207 full texts, and included 172 studies. The synthesis covers six themes: watermark embedding paradigms; threat models and attacker assumptions; evaluation metrics and experimental design; robustness under post-generation transformations; deployment readiness and access assumptions; and the limited evidence on application, governance, and misuse contexts. The field is methodologically rich but evidentially fragmented: many studies report strong detection performance while relying on narrow datasets, inconsistent metrics, limited baselines, implicit threat models, unrealistic transformations, or access assumptions unsuitable for closed or API-based deployments. Recurring concerns include false-positive control, public verification, adaptive paraphrasing, mixed-origin documents, benchmark fragmentation, and reproducibility. We therefore provide a descriptive evidence-completeness assessment and a research agenda for more robust, transparent, and deployment-ready LLM watermarking systems.
No takes yet. Share an insight, caveat, or question.
Islam et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: