Computational study demonstrates certified robustness across four word-level adversarial operations in language models, indicating improved defense against text-based attacks.
Key Points
Text-CRS establishes certified robustness against four word-level adversarial operations in classification models, achieving significant accuracy improvements over previous methods.
Theoretical modeling with randomized smoothing evaluates word alterations by mapping perturbations into combined permutation and embedding transformation spaces.
Highlights a generalized certification benchmark for diverse text attacks, outperforming state-of-the-art defenses against synonym substitution across multiple language models.