Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
Fuente:
arXiv
Saved in:
| Main Authors: | Yoon, Sangyeon, Hong, Hyesoo, Jeung, Wonje, No, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
by: Hong, Hyesoo, et al.
Published: (2026)
by: Hong, Hyesoo, et al.
Published: (2026)
Adversarial Sample-Based Approach for Tighter Privacy Auditing in Final Model-Only Scenarios
by: Yoon, Sangyeon, et al.
Published: (2024)
by: Yoon, Sangyeon, et al.
Published: (2024)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
An Information Theoretic Evaluation Metric For Strong Unlearning
by: Jeon, Dongjae, et al.
Published: (2024)
by: Jeon, Dongjae, et al.
Published: (2024)
A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
R-TOFU: Unlearning in Large Reasoning Models
by: Yoon, Sangyeon, et al.
Published: (2025)
by: Yoon, Sangyeon, et al.
Published: (2025)
SEPS: A Separability Measure for Robust Unlearning in LLMs
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
DUSK: Do Not Unlearn Shared Knowledge
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
by: Ha, SeungBum, et al.
Published: (2025)
by: Ha, SeungBum, et al.
Published: (2025)
FedCARE: Federated Unlearning with Conflict-Aware Projection and Relearning-Resistant Recovery
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024)
by: Shin, Jeongjin, et al.
Published: (2024)
Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
by: Hu, Shengyuan, et al.
Published: (2024)
by: Hu, Shengyuan, et al.
Published: (2024)
Multi-Level Knowledge Distillation and Dynamic Self-Supervised Learning for Continual Learning
by: Kim, Taeheon, et al.
Published: (2025)
by: Kim, Taeheon, et al.
Published: (2025)
NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Federated Unlearning in the Wild: Rethinking Fairness and Data Discrepancy
by: Huang, ZiHeng, et al.
Published: (2025)
by: Huang, ZiHeng, et al.
Published: (2025)
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
by: Polowczyk, Agnieszka, et al.
Published: (2025)
by: Polowczyk, Agnieszka, et al.
Published: (2025)
Forgetting is Competition: Rethinking Unlearning as Representation Interference in Diffusion Models
by: Ranjan, Ashutosh, et al.
Published: (2026)
by: Ranjan, Ashutosh, et al.
Published: (2026)
Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
by: Kim, Bumjun, et al.
Published: (2025)
by: Kim, Bumjun, et al.
Published: (2025)
Layered Unlearning for Adversarial Relearning
by: Qian, Timothy, et al.
Published: (2025)
by: Qian, Timothy, et al.
Published: (2025)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Activation by Interval-wise Dropout: A Simple Way to Prevent Neural Networks from Plasticity Loss
by: Park, Sangyeon, et al.
Published: (2025)
by: Park, Sangyeon, et al.
Published: (2025)
Incremental Learning of Retrievable Skills For Efficient Continual Task Adaptation
by: Lee, Daehee, et al.
Published: (2024)
by: Lee, Daehee, et al.
Published: (2024)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
by: Sun, Shaoning, et al.
Published: (2026)
by: Sun, Shaoning, et al.
Published: (2026)
Learning Equi-angular Representations for Online Continual Learning
by: Seo, Minhyuk, et al.
Published: (2024)
by: Seo, Minhyuk, et al.
Published: (2024)
Benign Overfitting in Adversarial Training for Vision Transformers
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Deep Unlearn: Benchmarking Machine Unlearning for Image Classification
by: Cadet, Xavier F., et al.
Published: (2024)
by: Cadet, Xavier F., et al.
Published: (2024)
Rethinking Forward Processes for Score-Based Nonlinear Data Assimilation in High Dimensions
by: Yoon, Eunbi, et al.
Published: (2026)
by: Yoon, Eunbi, et al.
Published: (2026)
Low Resource Reconstruction Attacks Through Benign Prompts
by: Yarkoni, Sol, et al.
Published: (2025)
by: Yarkoni, Sol, et al.
Published: (2025)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
by: Mahmud, Tamim Al, et al.
Published: (2025)
by: Mahmud, Tamim Al, et al.
Published: (2025)
Unlearning Information Bottleneck: Machine Unlearning of Systematic Patterns and Biases
by: Han, Ling, et al.
Published: (2024)
by: Han, Ling, et al.
Published: (2024)
On the Emergence of Syntax by Means of Local Interaction
by: Wei, Zichao
Published: (2026)
by: Wei, Zichao
Published: (2026)
Rethinking Post-Unlearning Behavior of Large Vision-Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
Generating Synthetic Fair Syntax-agnostic Data by Learning and Distilling Fair Representation
by: Sikder, Md Fahim, et al.
Published: (2024)
by: Sikder, Md Fahim, et al.
Published: (2024)
A Classical View on Benign Overfitting: The Role of Sample Size
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Distillation Robustifies Unlearning
by: Lee, Bruce W., et al.
Published: (2025)
by: Lee, Bruce W., et al.
Published: (2025)
Similar Items
-
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
by: Hong, Hyesoo, et al.
Published: (2026) -
Adversarial Sample-Based Approach for Tighter Privacy Auditing in Final Model-Only Scenarios
by: Yoon, Sangyeon, et al.
Published: (2024) -
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026) -
An Information Theoretic Evaluation Metric For Strong Unlearning
by: Jeon, Dongjae, et al.
Published: (2024) -
A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
by: Jeung, Wonje, et al.
Published: (2025)