Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
Fuente:
arXiv
Guardado en:
| Autores principales: | Khoriaty, Matthew, Shportko, Andrii, Mercier, Gustavo, Wood-Doughty, Zach |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Controlling for Unobserved Confounding with Large Language Model Classification of Patient Smoking Status
por: Lee, Samuel, et al.
Publicado: (2024)
por: Lee, Samuel, et al.
Publicado: (2024)
Don't Forget Imagination!
por: Vityaev, Evgenii E., et al.
Publicado: (2025)
por: Vityaev, Evgenii E., et al.
Publicado: (2025)
"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise
por: Khoriaty, Matthew, et al.
Publicado: (2026)
por: Khoriaty, Matthew, et al.
Publicado: (2026)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
por: Plyusov, Daniil, et al.
Publicado: (2026)
por: Plyusov, Daniil, et al.
Publicado: (2026)
Kolmogorov Complexity Bounds for LLM Steganography and a Perplexity-Based Detection Proxy
por: Shportko, Andrii
Publicado: (2026)
por: Shportko, Andrii
Publicado: (2026)
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
por: Poole, Benjamin, et al.
Publicado: (2026)
por: Poole, Benjamin, et al.
Publicado: (2026)
Don't Forget to Connect! Improving RAG with Graph-based Reranking
por: Dong, Jialin, et al.
Publicado: (2024)
por: Dong, Jialin, et al.
Publicado: (2024)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
por: Cywiński, Bartosz, et al.
Publicado: (2025)
por: Cywiński, Bartosz, et al.
Publicado: (2025)
Forget and Explain: Transparent Verification of GNN Unlearning
por: Ahsan, Imran, et al.
Publicado: (2025)
por: Ahsan, Imran, et al.
Publicado: (2025)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
por: Zhang, Jiefu, et al.
Publicado: (2026)
por: Zhang, Jiefu, et al.
Publicado: (2026)
Finding Belief Geometries with Sparse Autoencoders
por: Levinson, Matthew
Publicado: (2026)
por: Levinson, Matthew
Publicado: (2026)
Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
por: Balashov, Andrii
Publicado: (2025)
por: Balashov, Andrii
Publicado: (2025)
Paraphrasing Attack Resilience of Various AI-Generated Text Detection Methods
por: Shportko, Andrii, et al.
Publicado: (2026)
por: Shportko, Andrii, et al.
Publicado: (2026)
Transformers Don't In-Context Learn Least Squares Regression
por: Hill, Joshua, et al.
Publicado: (2025)
por: Hill, Joshua, et al.
Publicado: (2025)
Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models
por: Garg, Arpit, et al.
Publicado: (2025)
por: Garg, Arpit, et al.
Publicado: (2025)
Forgetting is Competition: Rethinking Unlearning as Representation Interference in Diffusion Models
por: Ranjan, Ashutosh, et al.
Publicado: (2026)
por: Ranjan, Ashutosh, et al.
Publicado: (2026)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
por: Muhamed, Aashiq, et al.
Publicado: (2025)
por: Muhamed, Aashiq, et al.
Publicado: (2025)
Unlearning via Sparse Representations
por: Shah, Vedant, et al.
Publicado: (2023)
por: Shah, Vedant, et al.
Publicado: (2023)
Sparse Autoencoders, Again?
por: Lu, Yin, et al.
Publicado: (2025)
por: Lu, Yin, et al.
Publicado: (2025)
Benchmarking is Broken -- Don't Let AI be its Own Judge
por: Cheng, Zerui, et al.
Publicado: (2025)
por: Cheng, Zerui, et al.
Publicado: (2025)
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
por: Kim, Hyunseung, et al.
Publicado: (2024)
por: Kim, Hyunseung, et al.
Publicado: (2024)
Don't Waste Your Time: Early Stopping Cross-Validation
por: Bergman, Edward, et al.
Publicado: (2024)
por: Bergman, Edward, et al.
Publicado: (2024)
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
por: Hu, Shengyuan, et al.
Publicado: (2025)
por: Hu, Shengyuan, et al.
Publicado: (2025)
OPC: One-Point-Contraction Unlearning Toward Deep Feature Forgetting
por: Jung, Jaeheun, et al.
Publicado: (2025)
por: Jung, Jaeheun, et al.
Publicado: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
por: Bhalla, Usha, et al.
Publicado: (2026)
por: Bhalla, Usha, et al.
Publicado: (2026)
Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
por: Kazoom, Roie, et al.
Publicado: (2025)
por: Kazoom, Roie, et al.
Publicado: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
por: Dey, Nolan, et al.
Publicado: (2025)
por: Dey, Nolan, et al.
Publicado: (2025)
xAI-Drop: Don't Use What You Cannot Explain
por: De Luca, Vincenzo Marco, et al.
Publicado: (2024)
por: De Luca, Vincenzo Marco, et al.
Publicado: (2024)
FaLW: A Forgetting-aware Loss Reweighting for Long-tailed Unlearning
por: Yu, Liheng, et al.
Publicado: (2026)
por: Yu, Liheng, et al.
Publicado: (2026)
Are Sparse Autoencoder Benchmarks Reliable?
por: Chanin, David
Publicado: (2026)
por: Chanin, David
Publicado: (2026)
Machine Unlearning using Forgetting Neural Networks
por: Hatua, Amartya, et al.
Publicado: (2024)
por: Hatua, Amartya, et al.
Publicado: (2024)
Distill, Forget, Repeat: A Framework for Continual Unlearning in Text-to-Image Diffusion Models
por: George, Naveen, et al.
Publicado: (2025)
por: George, Naveen, et al.
Publicado: (2025)
Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
por: Liu, Zhihao, et al.
Publicado: (2025)
por: Liu, Zhihao, et al.
Publicado: (2025)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
por: Wang, Changsheng, et al.
Publicado: (2025)
por: Wang, Changsheng, et al.
Publicado: (2025)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
por: Park, Young-Jin, et al.
Publicado: (2025)
por: Park, Young-Jin, et al.
Publicado: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
por: Bajpai, Divya Jyoti, et al.
Publicado: (2025)
por: Bajpai, Divya Jyoti, et al.
Publicado: (2025)
Don't throw the baby out with the bathwater: How and why deep learning for ARC
por: Cole, Jack, et al.
Publicado: (2025)
por: Cole, Jack, et al.
Publicado: (2025)
Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
por: Lade, Ankit Hemant, et al.
Publicado: (2026)
por: Lade, Ankit Hemant, et al.
Publicado: (2026)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
por: Yang, Yongyi, et al.
Publicado: (2026)
por: Yang, Yongyi, et al.
Publicado: (2026)
Ejemplares similares
-
Controlling for Unobserved Confounding with Large Language Model Classification of Patient Smoking Status
por: Lee, Samuel, et al.
Publicado: (2024) -
Don't Forget Imagination!
por: Vityaev, Evgenii E., et al.
Publicado: (2025) -
"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise
por: Khoriaty, Matthew, et al.
Publicado: (2026) -
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
por: Plyusov, Daniil, et al.
Publicado: (2026) -
Kolmogorov Complexity Bounds for LLM Steganography and a Perplexity-Based Detection Proxy
por: Shportko, Andrii
Publicado: (2026)