Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shilov, Igor, Cloud, Alex, Gema, Aryo Pradipta, Goldman-Wetzler, Jacob, Panickssery, Nina, Sleight, Henry, Jones, Erik, Anil, Cem |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
by: Hägele, Alexander, et al.
Published: (2026)
by: Hägele, Alexander, et al.
Published: (2026)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024)
by: Cloud, Alex, et al.
Published: (2024)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025)
by: Gema, Aryo Pradipta, et al.
Published: (2025)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
by: Rajani, Neel, et al.
Published: (2025)
by: Rajani, Neel, et al.
Published: (2025)
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023)
by: Gema, Aryo Pradipta, et al.
Published: (2023)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Distillation Robustifies Unlearning
by: Lee, Bruce W., et al.
Published: (2025)
by: Lee, Bruce W., et al.
Published: (2025)
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
by: Kwan, Wai-Chung, et al.
Published: (2026)
by: Kwan, Wai-Chung, et al.
Published: (2026)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
by: Ackerman, Christopher, et al.
Published: (2024)
by: Ackerman, Christopher, et al.
Published: (2024)
Speeding up and reducing memory usage for scientific machine learning via mixed precision
by: Hayford, Joel, et al.
Published: (2024)
by: Hayford, Joel, et al.
Published: (2024)
Mitigating Many-Shot Jailbreaking
by: Ackerman, Christopher M., et al.
Published: (2025)
by: Ackerman, Christopher M., et al.
Published: (2025)
Self-Training Large Language Models for Tool-Use Without Demonstrations
by: Luo, Ne, et al.
Published: (2025)
by: Luo, Ne, et al.
Published: (2025)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
GRADA: Graph-based Reranking against Adversarial Documents Attack
by: Zheng, Jingjie, et al.
Published: (2025)
by: Zheng, Jingjie, et al.
Published: (2025)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024)
by: Ball, Sarah, et al.
Published: (2024)
Can GPT-3.5 Generate and Code Discharge Summaries?
by: Falis, Matúš, et al.
Published: (2024)
by: Falis, Matúš, et al.
Published: (2024)
A Comparative Study on Patient Language across Therapeutic Domains for Effective Patient Voice Classification in Online Health Discussions
by: Lysandrou, Giorgos, et al.
Published: (2024)
by: Lysandrou, Giorgos, et al.
Published: (2024)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
by: Murphy, Alexander, et al.
Published: (2025)
by: Murphy, Alexander, et al.
Published: (2025)
Unsupervised Elicitation of Language Models
by: Wen, Jiaxin, et al.
Published: (2025)
by: Wen, Jiaxin, et al.
Published: (2025)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
by: Højmark, Axel, et al.
Published: (2024)
by: Højmark, Axel, et al.
Published: (2024)
Abstractive Red-Teaming of Language Model Character
by: Rahn, Nate, et al.
Published: (2026)
by: Rahn, Nate, et al.
Published: (2026)
Feminist mental health activism in England c. 1968–95. By K.Mahoney, Manchester: Manchester University Press. 2023. pp. 280. £85.00 (cloth), £85.00 (ebk). ISBN: 9781526162267
by: Sara Wetzler
Published: (2024)
by: Sara Wetzler
Published: (2024)
Fight Books in Comparative Perspective. An Introduction
by: Sixt Wetzler
Published: (2020)
by: Sixt Wetzler
Published: (2020)
Paul Bowman, Mythologies of Martial Arts (Martial Arts Studies, 2), London/New York: Rowman and Littlefield, 2017, xxiii + 186 p.
by: Sixt Wetzler
Published: (2017)
by: Sixt Wetzler
Published: (2017)
“Your Kung Fu is very good, Master Fiore!” Asian and European fight books in comparison
by: Sixt Wetzler
Published: (2016)
by: Sixt Wetzler
Published: (2016)
PERSPECTIVES ON URBAN LONELINESS : City Paths Toward (Un)common Lifeworlds
by: Nina Goldman
Published: (2026)
by: Nina Goldman
Published: (2026)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
No Difference in Pullout Strength Between a Bio‐inductive Implant and a Semitendinosus Tendon Graft in a Biomechanical Study of Medial Patellofemoral Ligament Repair Augmentation
by: Austin Wetzler, et al.
Published: (2024)
by: Austin Wetzler, et al.
Published: (2024)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
by: Guo, Shiyuan, et al.
Published: (2025)
by: Guo, Shiyuan, et al.
Published: (2025)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
by: Ensign, Danielle, et al.
Published: (2025)
by: Ensign, Danielle, et al.
Published: (2025)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
by: Jones, Erik, et al.
Published: (2025)
by: Jones, Erik, et al.
Published: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
by: Wani, Farooq Ahmad, et al.
Published: (2026)
by: Wani, Farooq Ahmad, et al.
Published: (2026)
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
by: Rinberg, Roy, et al.
Published: (2025)
by: Rinberg, Roy, et al.
Published: (2025)
Similar Items
-
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
by: Hägele, Alexander, et al.
Published: (2026) -
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024) -
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025) -
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025) -
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
by: Rajani, Neel, et al.
Published: (2025)