Abstract
Large language models trained on vast corpora inherently risk memorizing sensitive or harmful content, which may later resurface in their outputs. Prevailing unlearning methods generally rely on gradient ascent and its variants to lower the probability of specific target responses. However, we find that this strategy induces a critical side effect: probability mass is redistributed into high-likelihood regions, often corresponding to semantically related rephrasings of the targets. We refer to this as the squeezing effect, which explains why many methods yield merely spurious unlearning, a problem further obscured by automated metrics (e.g., ROUGE, truth ratio) that misreport actual success. To address this, we propose a bootstrapping (BS) framework that explicitly links the squeezing effect with the model’s own high-confidence generations, namely its model beliefs. Since model beliefs inherently capture the very high-likelihood regions where probability mass is squeezed, incorporating them into the unlearning objective directly counters the squeezing effect. By jointly suppressing both target responses and model beliefs, BS-T (token) attenuates high-probability tokens, whereas BS-S (sequence) removes entire high-confidence generations, together achieving more thorough forgetting while preserving utility. Extensive experiments on diverse benchmarks confirm the effectiveness of our approach, with code merged to OpenUnlearning.
| Original language | English |
|---|---|
| Title of host publication | The Fourteenth International Conference on Learning Representations, ICLR 2026 |
| Publisher | International Conference on Learning Representations, ICLR |
| Pages | 1-36 |
| Number of pages | 36 |
| Publication status | Published - 23 Apr 2026 |
| Event | The Fourteenth International Conference on Learning Representations - Rio de Janeiro, Brazil, Rio de Janeirol, Brazil Duration: 23 Apr 2026 → 23 Apr 2026 https://openreview.net/group?id=ICLR.cc/2026/Conference#tab-accept-oral |
Publication series
| Name | International Conference on Learning Representations, ICLR |
|---|
Conference
| Conference | The Fourteenth International Conference on Learning Representations |
|---|---|
| Abbreviated title | ICLR 2026 |
| Country/Territory | Brazil |
| City | Rio de Janeirol |
| Period | 23/04/26 → 23/04/26 |
| Internet address |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 16 Peace, Justice and Strong Institutions
User-Defined Keywords
- Large Language Model Unlearning
Fingerprint
Dive into the research topics of 'LLM Unlearning with LLM Beliefs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver