CRISP: Persistent Concept Unlearning via Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Ashuach, Tomer, Arad, Dana, Mueller, Aaron, Tutek, Martin, Belinkov, Yonatan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
by: Ashuach, Tomer, et al.
Published: (2024)
by: Ashuach, Tomer, et al.
Published: (2024)
Reasoning Models Know What's Important, and Encode It in Their Activations
by: Nikankin, Yaniv, et al.
Published: (2026)
by: Nikankin, Yaniv, et al.
Published: (2026)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
by: Ashuach, Tomer, et al.
Published: (2026)
by: Ashuach, Tomer, et al.
Published: (2026)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
by: Nikankin, Yaniv, et al.
Published: (2025)
by: Nikankin, Yaniv, et al.
Published: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
by: Simhi, Adi, et al.
Published: (2026)
by: Simhi, Adi, et al.
Published: (2026)
HACK: Hallucinations Along Certainty and Knowledge Axes
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Are formal and functional linguistic mechanisms dissociated in language models?
by: Hanna, Michael, et al.
Published: (2025)
by: Hanna, Michael, et al.
Published: (2025)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
by: Nikankin, Yaniv, et al.
Published: (2024)
by: Nikankin, Yaniv, et al.
Published: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
Distinguishing Ignorance from Error in LLM Hallucinations
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
by: Hanna, Michael, et al.
Published: (2024)
by: Hanna, Michael, et al.
Published: (2024)
Position-aware Automatic Circuit Discovery
by: Haklay, Tal, et al.
Published: (2025)
by: Haklay, Tal, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models
by: Lee, Hwiyeong, et al.
Published: (2025)
by: Lee, Hwiyeong, et al.
Published: (2025)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
by: Stacey, Joe, et al.
Published: (2025)
by: Stacey, Joe, et al.
Published: (2025)
Combining Denoising Autoencoders with Contrastive Learning to fine-tune Transformer Models
by: Lopez-Avila, Alejo, et al.
Published: (2024)
by: Lopez-Avila, Alejo, et al.
Published: (2024)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
by: Rudman, William, et al.
Published: (2026)
by: Rudman, William, et al.
Published: (2026)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Parsing Akkadian Verbs with Prolog
by: Macks, Aaron
Published: (2024)
by: Macks, Aaron
Published: (2024)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
by: Patel, Het, et al.
Published: (2026)
by: Patel, Het, et al.
Published: (2026)
Graphemic Normalization of the Perso-Arabic Script
by: Doctor, Raiomond, et al.
Published: (2022)
by: Doctor, Raiomond, et al.
Published: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023)
by: Gutkin, Alexander, et al.
Published: (2023)
Comparing Complex Concepts with Transformers: Matching Patent Claims Against Natural Language Text
by: Blume, Matthias, et al.
Published: (2024)
by: Blume, Matthias, et al.
Published: (2024)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
by: Abitante, João Vitor Boer, et al.
Published: (2026)
by: Abitante, João Vitor Boer, et al.
Published: (2026)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
by: Nzeyimana, Antoine, et al.
Published: (2025)
by: Nzeyimana, Antoine, et al.
Published: (2025)
Learning Translations via Matrix Completion
by: Wijaya, Derry, et al.
Published: (2024)
by: Wijaya, Derry, et al.
Published: (2024)
Low-resource neural machine translation with morphological modeling
by: Nzeyimana, Antoine
Published: (2024)
by: Nzeyimana, Antoine
Published: (2024)
Efficient Reasoning via Thought-Training and Thought-Free Inference
by: Wu, Canhui, et al.
Published: (2025)
by: Wu, Canhui, et al.
Published: (2025)
ToolGen: Unified Tool Retrieval and Calling via Generation
by: Wang, Renxi, et al.
Published: (2024)
by: Wang, Renxi, et al.
Published: (2024)
Improving Retrospective Language Agents via Joint Policy Gradient Optimization
by: Feng, Xueyang, et al.
Published: (2025)
by: Feng, Xueyang, et al.
Published: (2025)
Similar Items
-
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
by: Ashuach, Tomer, et al.
Published: (2024) -
Reasoning Models Know What's Important, and Encode It in Their Activations
by: Nikankin, Yaniv, et al.
Published: (2026) -
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
by: Ashuach, Tomer, et al.
Published: (2026) -
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025) -
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
by: Nikankin, Yaniv, et al.
Published: (2025)