LLM Unlearning via Neural Activation Redirection
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, William F., Qiu, Xinchi, Kurmanji, Meghdad, Iacob, Alex, Sani, Lorenzo, Chen, Yihong, Cancedda, Nicola, Lane, Nicholas D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
by: Iacob, Alex, et al.
Published: (2025)
by: Iacob, Alex, et al.
Published: (2025)
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
by: Aleksandrov, Preslav, et al.
Published: (2025)
by: Aleksandrov, Preslav, et al.
Published: (2025)
DEPT: Decoupled Embeddings for Pre-training Language Models
by: Iacob, Alex, et al.
Published: (2024)
by: Iacob, Alex, et al.
Published: (2024)
Position: Bridge the Gaps between Machine Unlearning and AI Regulation
by: Marino, Bill, et al.
Published: (2025)
by: Marino, Bill, et al.
Published: (2025)
SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention
by: Shen, William F., et al.
Published: (2025)
by: Shen, William F., et al.
Published: (2025)
Editing as Unlearning: Are Knowledge Editing Methods Strong Baselines for Large Language Model Unlearning?
by: Li, Zexi, et al.
Published: (2025)
by: Li, Zexi, et al.
Published: (2025)
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2026)
by: Tastan, Nurbek, et al.
Published: (2026)
DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models
by: Iacob, Alex, et al.
Published: (2025)
by: Iacob, Alex, et al.
Published: (2025)
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
by: Jovanović, Andrej, et al.
Published: (2026)
by: Jovanović, Andrej, et al.
Published: (2026)
$f$-FUM: Federated Unlearning via min--max and $f$-divergence
by: Karimian, Radmehr, et al.
Published: (2026)
by: Karimian, Radmehr, et al.
Published: (2026)
Worldwide Federated Training of Language Models
by: Iacob, Alex, et al.
Published: (2024)
by: Iacob, Alex, et al.
Published: (2024)
The Future of Large Language Model Pre-training is Federated
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
Task-Centric Personalized Federated Fine-Tuning of Language Models
by: Talasso, Gabriel U., et al.
Published: (2026)
by: Talasso, Gabriel U., et al.
Published: (2026)
Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
by: Wannan, et al.
Published: (2025)
by: Wannan, et al.
Published: (2025)
Spectral Filters, Dark Signals, and Attention Sinks
by: Cancedda, Nicola
Published: (2024)
by: Cancedda, Nicola
Published: (2024)
Photon: Federated LLM Pre-Training
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
Sheaf HyperNetworks for Personalized Federated Learning
by: Nguyen, Bao, et al.
Published: (2024)
by: Nguyen, Bao, et al.
Published: (2024)
Pollen: High-throughput Federated Learning Simulation via Resource-Aware Client Placement
by: Sani, Lorenzo, et al.
Published: (2023)
by: Sani, Lorenzo, et al.
Published: (2023)
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
FedAnchor: Enhancing Federated Semi-Supervised Learning with Label Contrastive Loss for Unlabeled Clients
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
SparsyFed: Sparse Adaptive Federated Training
by: Guastella, Adriano, et al.
Published: (2025)
by: Guastella, Adriano, et al.
Published: (2025)
Gradient-less Federated Gradient Boosting Trees with Learnable Learning Rates
by: Ma, Chenyang, et al.
Published: (2023)
by: Ma, Chenyang, et al.
Published: (2023)
SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models
by: Liu, Renyang, et al.
Published: (2026)
by: Liu, Renyang, et al.
Published: (2026)
Inference-Time Machine Unlearning via Gated Activation Redirection
by: Turani, Vinícius Conte, et al.
Published: (2026)
by: Turani, Vinícius Conte, et al.
Published: (2026)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
by: Koishekenov, Yeskendir, et al.
Published: (2025)
by: Koishekenov, Yeskendir, et al.
Published: (2025)
Measuring the Depth of LLM Unlearning via Activation Patching
by: Lee, Jaeung, et al.
Published: (2026)
by: Lee, Jaeung, et al.
Published: (2026)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
Verifying Chain-of-Thought Reasoning via Its Computational Graph
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
Breaking Physical and Linguistic Borders: Multilingual Federated Prompt Tuning for Low-Resource Languages
by: Zhao, Wanru, et al.
Published: (2025)
by: Zhao, Wanru, et al.
Published: (2025)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
by: Shen, William F., et al.
Published: (2026)
by: Shen, William F., et al.
Published: (2026)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs
by: Pan, Chao, et al.
Published: (2026)
by: Pan, Chao, et al.
Published: (2026)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Unlink to Unlearn: Simplifying Edge Unlearning in GNNs
by: Tan, Jiajun, et al.
Published: (2024)
by: Tan, Jiajun, et al.
Published: (2024)
ContinualFlow: Learning and Unlearning with Neural Flow Matching
by: Simone, Lorenzo, et al.
Published: (2025)
by: Simone, Lorenzo, et al.
Published: (2025)
Evaluating Privacy Leakage in Split Learning
by: Qiu, Xinchi, et al.
Published: (2023)
by: Qiu, Xinchi, et al.
Published: (2023)
TRAP: Targeted Redirecting of Agentic Preferences
by: Kang, Hangoo, et al.
Published: (2025)
by: Kang, Hangoo, et al.
Published: (2025)
What makes unlearning hard and what to do about it
by: Zhao, Kairan, et al.
Published: (2024)
by: Zhao, Kairan, et al.
Published: (2024)
Catastrophic Failure of LLM Unlearning via Quantization
by: Zhang, Zhiwei, et al.
Published: (2024)
by: Zhang, Zhiwei, et al.
Published: (2024)
Similar Items
-
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024) -
MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
by: Iacob, Alex, et al.
Published: (2025) -
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
by: Aleksandrov, Preslav, et al.
Published: (2025) -
DEPT: Decoupled Embeddings for Pre-training Language Models
by: Iacob, Alex, et al.
Published: (2024) -
Position: Bridge the Gaps between Machine Unlearning and AI Regulation
by: Marino, Bill, et al.
Published: (2025)