Precise In-Parameter Concept Erasure in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gur-Arieh, Yoav, Suslik, Clara, Hong, Yihuai, Barez, Fazl, Geva, Mor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
by: Gur-Arieh, Yoav, et al.
Published: (2026)
by: Gur-Arieh, Yoav, et al.
Published: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
by: Avrahamy, Asaf, et al.
Published: (2026)
by: Avrahamy, Asaf, et al.
Published: (2026)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
by: Gottesman, Daniela, et al.
Published: (2025)
by: Gottesman, Daniela, et al.
Published: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Estimating Knowledge in Large Language Models Without Generating a Single Token
by: Gottesman, Daniela, et al.
Published: (2024)
by: Gottesman, Daniela, et al.
Published: (2024)
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)
by: Elhelo, Amit, et al.
Published: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
by: Hong, Yihuai, et al.
Published: (2024)
by: Hong, Yihuai, et al.
Published: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Detecting (Un)answerability in Large Language Models with Linear Directions
by: Lavi, Maor Juliet, et al.
Published: (2025)
by: Lavi, Maor Juliet, et al.
Published: (2025)
Kernelized Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Linear Adversarial Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
by: Yalon, Noam Steinmetz, et al.
Published: (2026)
by: Yalon, Noam Steinmetz, et al.
Published: (2026)
The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
by: Hong, Yihuai, et al.
Published: (2025)
by: Hong, Yihuai, et al.
Published: (2025)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
by: Yang, Sohee, et al.
Published: (2024)
by: Yang, Sohee, et al.
Published: (2024)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
by: Lan, Michael, et al.
Published: (2024)
by: Lan, Michael, et al.
Published: (2024)
Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
by: Biran, Eden, et al.
Published: (2024)
by: Biran, Eden, et al.
Published: (2024)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
by: Yang, Sohee, et al.
Published: (2024)
by: Yang, Sohee, et al.
Published: (2024)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
by: Fu, Tingchen, et al.
Published: (2024)
by: Fu, Tingchen, et al.
Published: (2024)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
by: Cohen, Ido, et al.
Published: (2024)
by: Cohen, Ido, et al.
Published: (2024)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Large Language Models Relearn Removed Concepts
by: Lo, Michelle, et al.
Published: (2024)
by: Lo, Michelle, et al.
Published: (2024)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
by: Yona, Itay, et al.
Published: (2026)
by: Yona, Itay, et al.
Published: (2026)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
by: Ivgi, Maor, et al.
Published: (2024)
by: Ivgi, Maor, et al.
Published: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Eliciting Textual Descriptions from Representations of Continuous Prompts
by: Ramati, Dana, et al.
Published: (2024)
by: Ramati, Dana, et al.
Published: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
by: Yona, Gal, et al.
Published: (2026)
by: Yona, Gal, et al.
Published: (2026)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
by: Shafran, Or, et al.
Published: (2026)
by: Shafran, Or, et al.
Published: (2026)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
Preventing Rogue Agents Improves Multi-Agent Collaboration
by: Barbi, Ohav, et al.
Published: (2025)
by: Barbi, Ohav, et al.
Published: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
by: Ahrac, Sagi, et al.
Published: (2026)
by: Ahrac, Sagi, et al.
Published: (2026)
Similar Items
-
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
by: Gur-Arieh, Yoav, et al.
Published: (2025) -
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
by: Gur-Arieh, Yoav, et al.
Published: (2026) -
Disentangling MLP Neuron Weights in Vocabulary Space
by: Avrahamy, Asaf, et al.
Published: (2026) -
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025) -
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
by: Gottesman, Daniela, et al.
Published: (2025)