Comprehensive analysis reveals critical vulnerabilities in LLM training and inference, emphasizing robust countermeasures.
Large language models (LLMs) have dramatically reshaped the field of natural language processing, presenting groundbreaking advancements in many areas, from chatbots to content creation. However, with the increasing adoption of these sophisticated models, it is crucial to scrutinize the vulnerabilities associated with their training and inference stages. This comprehensive analysis highlights the critical threats and inefficiencies inherent to these processes and emphasizes the need for robust countermeasures. This paper presents an extensive study of training and inference time vulnerabilities in Large Language Models (LLMs), specifically focusing on poisoning, backdoor, paraphrasing, and spoofing attacks. We introduce novel evaluation frameworks and detection mechanisms for each attack type. Our experimental results across multiple attack vectors demonstrate varying degrees of model susceptibility and reveal critical security implications. The proposed defensive mechanisms showcase impressive model performance, highlighted by consistent successful evaluation outcomes.
No takes yet. Share an insight, caveat, or question.
Canan Batur Şahin (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: