Key points are not available for this paper at this time.
ABSTRACT Machine unlearning for large language models (LLMs) aims to selectively remove target knowledge while preserving non‐target knowledge. Existing methods often overlook the semantic entanglement between knowledge to be forgotten and knowledge to be retained, which may degrade retention performance. Motivated by this observation, retain‐subset construction is reformulated as an entanglement‐aware knowledge selection problem, and a plug‐and‐play framework is proposed. Specifically, the proposed framework quantifies knowledge coupling through a bounded Entanglement Score, and combines retrieval‐based selection with diversity‐aware sampling to construct representative retain subsets that require stronger protection. The selected data can be directly incorporated into existing unlearning algorithms, such as NPO and SimNPO, without modifying their loss functions. Experiments on TOFU and MUSE demonstrate improved utility–forgetting trade‐offs and reduced privacy leakage, while results on WMDP indicate that the framework is particularly beneficial when the knowledge to be forgotten and the knowledge to be retained are strongly coupled. Overall, this framework provides a general and plug‐and‐play approach to enhancing knowledge preservation in LLM unlearning.
Xiong et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: