Large Language Models Relearn Removed Concepts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lo, Michelle, Cohen, Shay B., Barez, Fazl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
Query Circuits: Explaining How Language Models Answer User Prompts
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains
von: Fonseca, Marcio, et al.
Veröffentlicht: (2023)
von: Fonseca, Marcio, et al.
Veröffentlicht: (2023)
Rethinking AI Cultural Alignment
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
von: Fonseca, Marcio, et al.
Veröffentlicht: (2024)
von: Fonseca, Marcio, et al.
Veröffentlicht: (2024)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Visualizing Neural Network Imagination
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
Chain-of-Thought Hijacking
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
von: Zhou, Yi, et al.
Veröffentlicht: (2025)
von: Zhou, Yi, et al.
Veröffentlicht: (2025)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
von: Li, Changyi, et al.
Veröffentlicht: (2026)
von: Li, Changyi, et al.
Veröffentlicht: (2026)
Embodied AI: Emerging Risks and Opportunities for Policy Action
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
What can Large Language Models Capture about Code Functional Equivalence?
von: Maveli, Nickil, et al.
Veröffentlicht: (2024)
von: Maveli, Nickil, et al.
Veröffentlicht: (2024)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
von: Min, Nay Myat, et al.
Veröffentlicht: (2024)
von: Min, Nay Myat, et al.
Veröffentlicht: (2024)
Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour
von: Gyevnár, Bálint, et al.
Veröffentlicht: (2025)
von: Gyevnár, Bálint, et al.
Veröffentlicht: (2025)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
von: Denison, Carson, et al.
Veröffentlicht: (2024)
von: Denison, Carson, et al.
Veröffentlicht: (2024)
Precise In-Parameter Concept Erasure in Large Language Models
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
Spectral Editing of Activations for Large Language Model Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
`Keep it Together': Enforcing Cohesion in Extractive Summaries by Simulating Human Memory
von: Cardenas, Ronald, et al.
Veröffentlicht: (2024)
von: Cardenas, Ronald, et al.
Veröffentlicht: (2024)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language Models
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2026)
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2026)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
von: Deng, Ken, et al.
Veröffentlicht: (2026)
von: Deng, Ken, et al.
Veröffentlicht: (2026)
Modeling News Interactions and Influence for Financial Market Prediction
von: Wang, Mengyu, et al.
Veröffentlicht: (2024)
von: Wang, Mengyu, et al.
Veröffentlicht: (2024)
Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
von: Maveli, Nickil, et al.
Veröffentlicht: (2026)
von: Maveli, Nickil, et al.
Veröffentlicht: (2026)
FedCARE: Federated Unlearning with Conflict-Aware Projection and Relearning-Resistant Recovery
von: Li, Yue, et al.
Veröffentlicht: (2026)
von: Li, Yue, et al.
Veröffentlicht: (2026)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
UNLEARN Efficient Removal of Knowledge in Large Language Models
von: Lizzo, Tyler, et al.
Veröffentlicht: (2024)
von: Lizzo, Tyler, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024) -
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
von: Fu, Tingchen, et al.
Veröffentlicht: (2024) -
Query Circuits: Explaining How Language Models Answer User Prompts
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023) -
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)