Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Amara, Ibtihel, Humayun, Ahmed Imtiaz, Kajic, Ivana, Parekh, Zarana, Harris, Natalie, Young, Sarah, Nagpal, Chirag, Kim, Najoung, He, Junfeng, Vasconcelos, Cristina Nader, Ramachandran, Deepak, Farnadi, Golnoosh, Heller, Katherine, Havaei, Mohammad, Rostamzadeh, Negar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models
von: Humayun, Ahmed Imtiaz, et al.
Veröffentlicht: (2024)
von: Humayun, Ahmed Imtiaz, et al.
Veröffentlicht: (2024)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
von: Farnadi, Golnoosh, et al.
Veröffentlicht: (2024)
von: Farnadi, Golnoosh, et al.
Veröffentlicht: (2024)
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
von: Kassem, Aly, et al.
Veröffentlicht: (2026)
von: Kassem, Aly, et al.
Veröffentlicht: (2026)
Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models
von: Shi, Zhuan, et al.
Veröffentlicht: (2026)
von: Shi, Zhuan, et al.
Veröffentlicht: (2026)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
von: Braun, Tobias, et al.
Veröffentlicht: (2025)
von: Braun, Tobias, et al.
Veröffentlicht: (2025)
Erased or Dormant? Rethinking Concept Erasure Through Reversibility
von: Liu, Ping, et al.
Veröffentlicht: (2025)
von: Liu, Ping, et al.
Veröffentlicht: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
von: Farashah, Alireza Dehghanpour, et al.
Veröffentlicht: (2026)
von: Farashah, Alireza Dehghanpour, et al.
Veröffentlicht: (2026)
EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers
von: Gao, Daiheng, et al.
Veröffentlicht: (2024)
von: Gao, Daiheng, et al.
Veröffentlicht: (2024)
Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models
von: Seo, Hoigi, et al.
Veröffentlicht: (2026)
von: Seo, Hoigi, et al.
Veröffentlicht: (2026)
Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers
von: Jiang, Nanxiang, et al.
Veröffentlicht: (2026)
von: Jiang, Nanxiang, et al.
Veröffentlicht: (2026)
Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
von: Kusumba, Abhiram, et al.
Veröffentlicht: (2025)
von: Kusumba, Abhiram, et al.
Veröffentlicht: (2025)
FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models
von: Sun, Yi, et al.
Veröffentlicht: (2026)
von: Sun, Yi, et al.
Veröffentlicht: (2026)
LoRA Provides Differential Privacy by Design via Random Sketching
von: Malekmohammadi, Saber, et al.
Veröffentlicht: (2024)
von: Malekmohammadi, Saber, et al.
Veröffentlicht: (2024)
ActErase: A Training-Free Paradigm for Precise Concept Erasure via Activation Redirection
von: Sun, Yi, et al.
Veröffentlicht: (2026)
von: Sun, Yi, et al.
Veröffentlicht: (2026)
EraseAnything++: Enabling Concept Erasure in Rectified Flow Transformers Leveraging Multi-Object Optimization
von: Fan, Zhaoxin, et al.
Veröffentlicht: (2026)
von: Fan, Zhaoxin, et al.
Veröffentlicht: (2026)
The Case for Globalizing Fairness: A Mixed Methods Study on Colonialism, AI, and Health in Africa
von: Asiedu, Mercy, et al.
Veröffentlicht: (2024)
von: Asiedu, Mercy, et al.
Veröffentlicht: (2024)
Preference Models assume Proportional Hazards of Utilities
von: Nagpal, Chirag
Veröffentlicht: (2025)
von: Nagpal, Chirag
Veröffentlicht: (2025)
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
von: Das, Anudeep, et al.
Veröffentlicht: (2025)
von: Das, Anudeep, et al.
Veröffentlicht: (2025)
Mass Concept Erasure in Diffusion Models with Concept Hierarchy
von: Tu, Jiahang, et al.
Veröffentlicht: (2026)
von: Tu, Jiahang, et al.
Veröffentlicht: (2026)
Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
Fairness in Federated Learning: Fairness for Whom?
von: Taik, Afaf, et al.
Veröffentlicht: (2025)
von: Taik, Afaf, et al.
Veröffentlicht: (2025)
Multilingual Hallucination Gaps in Large Language Models
von: Chataigner, Cléa, et al.
Veröffentlicht: (2024)
von: Chataigner, Cléa, et al.
Veröffentlicht: (2024)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
von: Arzaghi, Mina, et al.
Veröffentlicht: (2024)
von: Arzaghi, Mina, et al.
Veröffentlicht: (2024)
Promoting Fair Vaccination Strategies Through Influence Maximization: A Case Study on COVID-19 Spread
von: Neophytou, Nicola, et al.
Veröffentlicht: (2024)
von: Neophytou, Nicola, et al.
Veröffentlicht: (2024)
Advancing Cultural Inclusivity: Optimizing Embedding Spaces for Balanced Music Recommendations
von: Moradi, Armin, et al.
Veröffentlicht: (2024)
von: Moradi, Armin, et al.
Veröffentlicht: (2024)
Differentially Private Clustered Federated Learning
von: Malekmohammadi, Saber, et al.
Veröffentlicht: (2024)
von: Malekmohammadi, Saber, et al.
Veröffentlicht: (2024)
Crossing Boundaries: Leveraging Semantic Divergences to Explore Cultural Novelty in Cooking Recipes
von: Carichon, Florian, et al.
Veröffentlicht: (2025)
von: Carichon, Florian, et al.
Veröffentlicht: (2025)
Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2026)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2026)
Data as a Lever: A Neighbouring Datasets Perspective on Predictive Multiplicity
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
von: More, Yash, et al.
Veröffentlicht: (2024)
von: More, Yash, et al.
Veröffentlicht: (2024)
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Erasing Concepts, Steering Generations: A Comprehensive Survey of Concept Suppression
von: Xie, Yiwei, et al.
Veröffentlicht: (2025)
von: Xie, Yiwei, et al.
Veröffentlicht: (2025)
Pruning for Robust Concept Erasing in Diffusion Models
von: Yang, Tianyun, et al.
Veröffentlicht: (2024)
von: Yang, Tianyun, et al.
Veröffentlicht: (2024)
ESC: Erasing Space Concept for Knowledge Deletion
von: Lee, Tae-Young, et al.
Veröffentlicht: (2025)
von: Lee, Tae-Young, et al.
Veröffentlicht: (2025)
When Are Concepts Erased From Diffusion Models?
von: Lu, Kevin, et al.
Veröffentlicht: (2025)
von: Lu, Kevin, et al.
Veröffentlicht: (2025)
Orthogonal Concept Erasure for Diffusion Models
von: Sun, Yuhao, et al.
Veröffentlicht: (2026)
von: Sun, Yuhao, et al.
Veröffentlicht: (2026)
Minimalist Concept Erasure in Generative Models
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models
von: Humayun, Ahmed Imtiaz, et al.
Veröffentlicht: (2024) -
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
von: Farnadi, Golnoosh, et al.
Veröffentlicht: (2024) -
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
von: Kassem, Aly, et al.
Veröffentlicht: (2026) -
Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models
von: Shi, Zhuan, et al.
Veröffentlicht: (2026) -
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)