Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Chengcan, Wei, Zeming, Chen, Huanran, Dong, Yinpeng, Sun, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
by: Zhou, Xueyang, et al.
Published: (2025)
by: Zhou, Xueyang, et al.
Published: (2025)
Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems
by: Parris, William
Published: (2026)
by: Parris, William
Published: (2026)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025)
by: Yamchote, Phaphontee, et al.
Published: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026)
by: Pepe, Alberto, et al.
Published: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
by: Tang, Zhengzheng
Published: (2026)
by: Tang, Zhengzheng
Published: (2026)
A Machine Learning Framework for Turbofan Health Estimation via Inverse Problem Formulation
by: Leyli-Abadi, Milad, et al.
Published: (2026)
by: Leyli-Abadi, Milad, et al.
Published: (2026)
HGTUL: A Hypergraph-based Model For Trajectory User Linking
by: Chang, Fengjie, et al.
Published: (2025)
by: Chang, Fengjie, et al.
Published: (2025)
HyperMask: Adaptive Hypernetwork-based Masks for Continual Learning
by: Książek, Kamil, et al.
Published: (2023)
by: Książek, Kamil, et al.
Published: (2023)
Is ReLU Adversarially Robust?
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Upside Down Reinforcement Learning with Policy Generators
by: Di Ventura, Jacopo, et al.
Published: (2025)
by: Di Ventura, Jacopo, et al.
Published: (2025)
Multi-Level Fusion Graph Neural Network for Molecule Property Prediction
by: Liu, XiaYu, et al.
Published: (2025)
by: Liu, XiaYu, et al.
Published: (2025)
CGLearn: Consistent Gradient-Based Learning for Out-of-Distribution Generalization
by: Chowdhury, Jawad, et al.
Published: (2024)
by: Chowdhury, Jawad, et al.
Published: (2024)
Concept Prerequisite Relation Prediction by Using Permutation-Equivariant Directed Graph Neural Networks
by: Qu, Xiran, et al.
Published: (2023)
by: Qu, Xiran, et al.
Published: (2023)
Simple Network Graph Comparative Learning
by: Yu, Qiang, et al.
Published: (2026)
by: Yu, Qiang, et al.
Published: (2026)
Mapping representations in Reinforcement Learning via Semantic Alignment for Zero-Shot Stitching
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
by: Turan, Berkant, et al.
Published: (2025)
by: Turan, Berkant, et al.
Published: (2025)
Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning
by: Filus, Katarzyna, et al.
Published: (2026)
by: Filus, Katarzyna, et al.
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization
by: Zhou, Xueyang, et al.
Published: (2025)
by: Zhou, Xueyang, et al.
Published: (2025)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
Batch Matrix-form Equations and Implementation of Multilayer Perceptrons
by: Wesselink, Wieger, et al.
Published: (2025)
by: Wesselink, Wieger, et al.
Published: (2025)
Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
by: Moore, Alexander, et al.
Published: (2025)
by: Moore, Alexander, et al.
Published: (2025)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
by: Keeman, Michael
Published: (2026)
by: Keeman, Michael
Published: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
by: Henry, James
Published: (2026)
by: Henry, James
Published: (2026)
A Practical Guide to Streaming Continual Learning
by: Cossu, Andrea, et al.
Published: (2026)
by: Cossu, Andrea, et al.
Published: (2026)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
by: Mathew, Aby Mammen
Published: (2026)
by: Mathew, Aby Mammen
Published: (2026)
cPNN: Continuous Progressive Neural Networks for Evolving Streaming Time Series
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
On the Origin of Algorithmic Progress in AI
by: Gundlach, Hans, et al.
Published: (2025)
by: Gundlach, Hans, et al.
Published: (2025)
ProfileXAI: User-Adaptive Explainable AI
by: Corrales, Gilber A., et al.
Published: (2025)
by: Corrales, Gilber A., et al.
Published: (2025)
Geometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
by: Loza, Andrew J., et al.
Published: (2025)
by: Loza, Andrew J., et al.
Published: (2025)
Decentralized Time Series Classification with ROCKET Features
by: Casella, Bruno, et al.
Published: (2025)
by: Casella, Bruno, et al.
Published: (2025)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
by: Ged, François, et al.
Published: (2023)
by: Ged, François, et al.
Published: (2023)
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
by: Henry, James
Published: (2026)
by: Henry, James
Published: (2026)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Similar Items
-
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
by: Zhou, Xueyang, et al.
Published: (2025) -
Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems
by: Parris, William
Published: (2026) -
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025) -
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026) -
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)