Nonlinear Concept Erasure: a Density Matching Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saillenfest, Antoine, Lemberger, Pirmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explaining Text Classifiers with Counterfactual Representations
von: Lemberger, Pirmin, et al.
Veröffentlicht: (2024)
von: Lemberger, Pirmin, et al.
Veröffentlicht: (2024)
How Graph Structure and Label Dependencies Contribute to Node Classification in a Large Network of Documents
von: Lemberger, Pirmin, et al.
Veröffentlicht: (2023)
von: Lemberger, Pirmin, et al.
Veröffentlicht: (2023)
ELIXIR: Efficient and LIghtweight model for eXplaIning Recommendations
von: Kabongo, Ben, et al.
Veröffentlicht: (2025)
von: Kabongo, Ben, et al.
Veröffentlicht: (2025)
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Revisiting Hierarchical Text Classification: Inference and Metrics
von: Plaud, Roman, et al.
Veröffentlicht: (2024)
von: Plaud, Roman, et al.
Veröffentlicht: (2024)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
TaCo: Targeted Concept Erasure Prevents Non-Linear Classifiers From Detecting Protected Attributes
von: Jourdan, Fanny, et al.
Veröffentlicht: (2023)
von: Jourdan, Fanny, et al.
Veröffentlicht: (2023)
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
von: Xue, Yuyang, et al.
Veröffentlicht: (2025)
von: Xue, Yuyang, et al.
Veröffentlicht: (2025)
Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
von: Feucht, Sheridan, et al.
Veröffentlicht: (2024)
von: Feucht, Sheridan, et al.
Veröffentlicht: (2024)
Obliviator Reveals the Cost of Nonlinear Guardedness in Concept Erasure
von: Akbari, Ramin, et al.
Veröffentlicht: (2026)
von: Akbari, Ramin, et al.
Veröffentlicht: (2026)
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
von: AlKhamissi, Badr, et al.
Veröffentlicht: (2024)
von: AlKhamissi, Badr, et al.
Veröffentlicht: (2024)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2024)
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2024)
Minimalist Concept Erasure in Generative Models
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
von: Yu, Xuemin, et al.
Veröffentlicht: (2026)
von: Yu, Xuemin, et al.
Veröffentlicht: (2026)
LLM Pretraining with Continuous Concepts
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
Towards Compositionality in Concept Learning
von: Stein, Adam, et al.
Veröffentlicht: (2024)
von: Stein, Adam, et al.
Veröffentlicht: (2024)
Medical Concept Normalization in a Low-Resource Setting
von: Patzelt, Tim
Veröffentlicht: (2024)
von: Patzelt, Tim
Veröffentlicht: (2024)
Learning Machines: In Search of a Concept Oriented Language
von: Gunes, Veyis
Veröffentlicht: (2024)
von: Gunes, Veyis
Veröffentlicht: (2024)
Concept Bottleneck Large Language Models
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
MatchXML: An Efficient Text-label Matching Framework for Extreme Multi-label Text Classification
von: Ye, Hui, et al.
Veröffentlicht: (2023)
von: Ye, Hui, et al.
Veröffentlicht: (2023)
Leverage Unlearning to Sanitize LLMs
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
AlignSAE: Concept-Aligned Sparse Autoencoders
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
Simple Mechanisms for Representing, Indexing and Manipulating Concepts
von: Li, Yuanzhi, et al.
Veröffentlicht: (2023)
von: Li, Yuanzhi, et al.
Veröffentlicht: (2023)
CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer
von: Wilam, Piotr
Veröffentlicht: (2026)
von: Wilam, Piotr
Veröffentlicht: (2026)
Khattat: Enhancing Readability and Concept Representation of Semantic Typography
von: Hussein, Ahmed, et al.
Veröffentlicht: (2024)
von: Hussein, Ahmed, et al.
Veröffentlicht: (2024)
Causality $\neq$ Invariance: Function and Concept Vectors in LLMs
von: Opiełka, Gustaw, et al.
Veröffentlicht: (2026)
von: Opiełka, Gustaw, et al.
Veröffentlicht: (2026)
Forecasting Events in Soccer Matches Through Language
von: Mendes-Neves, Tiago, et al.
Veröffentlicht: (2024)
von: Mendes-Neves, Tiago, et al.
Veröffentlicht: (2024)
Entity Matching using Large Language Models
von: Peeters, Ralph, et al.
Veröffentlicht: (2023)
von: Peeters, Ralph, et al.
Veröffentlicht: (2023)
Merging by Matching Models in Task Parameter Subspaces
von: Tam, Derek, et al.
Veröffentlicht: (2023)
von: Tam, Derek, et al.
Veröffentlicht: (2023)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
von: Nakkiran, Preetum, et al.
Veröffentlicht: (2025)
von: Nakkiran, Preetum, et al.
Veröffentlicht: (2025)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
von: Yan, Xinyuan, et al.
Veröffentlicht: (2025)
von: Yan, Xinyuan, et al.
Veröffentlicht: (2025)
Concept Algebra for (Score-Based) Text-Controlled Generative Models
von: Wang, Zihao, et al.
Veröffentlicht: (2023)
von: Wang, Zihao, et al.
Veröffentlicht: (2023)
Can LLMs Learn New Concepts Incrementally without Forgetting?
von: Zheng, Junhao, et al.
Veröffentlicht: (2024)
von: Zheng, Junhao, et al.
Veröffentlicht: (2024)
CLUE: Concept-Level Uncertainty Estimation for Large Language Models
von: Wang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
von: Wang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
Separable Multi-Concept Erasure from Diffusion Models
von: Zhao, Mengnan, et al.
Veröffentlicht: (2024)
von: Zhao, Mengnan, et al.
Veröffentlicht: (2024)
Adversarial Moment-Matching Distillation of Large Language Models
von: Jia, Chen
Veröffentlicht: (2024)
von: Jia, Chen
Veröffentlicht: (2024)
Ähnliche Einträge
-
Explaining Text Classifiers with Counterfactual Representations
von: Lemberger, Pirmin, et al.
Veröffentlicht: (2024) -
How Graph Structure and Label Dependencies Contribute to Node Classification in a Large Network of Documents
von: Lemberger, Pirmin, et al.
Veröffentlicht: (2023) -
ELIXIR: Efficient and LIghtweight model for eXplaIning Recommendations
von: Kabongo, Ben, et al.
Veröffentlicht: (2025) -
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022) -
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)