TaCo: Targeted Concept Erasure Prevents Non-Linear Classifiers From Detecting Protected Attributes
Fuente:
arXiv
Saved in:
| Main Authors: | Jourdan, Fanny, Béthune, Louis, Picard, Agustin, Risser, Laurent, Asher, Nicholas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability
by: Poché, Antonin, et al.
Published: (2025)
by: Poché, Antonin, et al.
Published: (2025)
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
by: Jourdan, Fanny
Published: (2024)
by: Jourdan, Fanny
Published: (2024)
Linear Adversarial Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes
by: Upadhayay, Bibek, et al.
Published: (2023)
by: Upadhayay, Bibek, et al.
Published: (2023)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
by: Fan, Yu, et al.
Published: (2025)
by: Fan, Yu, et al.
Published: (2025)
Kernelized Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Iterative Multilingual Spectral Attribute Erasure
by: Shao, Shun, et al.
Published: (2025)
by: Shao, Shun, et al.
Published: (2025)
TaCo: Capturing Spatio-Temporal Semantic Consistency in Remote Sensing Change Detection
by: Guo, Han, et al.
Published: (2025)
by: Guo, Han, et al.
Published: (2025)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
by: Jourdan, Fanny, et al.
Published: (2025)
by: Jourdan, Fanny, et al.
Published: (2025)
Extended Kohler's Rule of Magnetoresistance in TaCo$_2$Te$_2$
by: Pate, Samuel, et al.
Published: (2023)
by: Pate, Samuel, et al.
Published: (2023)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile Data
by: Cheng, Zhengxue, et al.
Published: (2026)
by: Cheng, Zhengxue, et al.
Published: (2026)
Precise In-Parameter Concept Erasure in Large Language Models
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
by: Bernas, Raphael, et al.
Published: (2026)
by: Bernas, Raphael, et al.
Published: (2026)
NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers
by: Greco, Salvatore, et al.
Published: (2024)
by: Greco, Salvatore, et al.
Published: (2024)
Nonlinear Concept Erasure: a Density Matching Approach
by: Saillenfest, Antoine, et al.
Published: (2025)
by: Saillenfest, Antoine, et al.
Published: (2025)
Magnetic field-induced non-trivial Lifshitz transition in TaCo2Te2
by: Pradhan, Suman Kalyan, et al.
Published: (2026)
by: Pradhan, Suman Kalyan, et al.
Published: (2026)
Interpreto: An Explainability Library for Transformers
by: Poché, Antonin, et al.
Published: (2025)
by: Poché, Antonin, et al.
Published: (2025)
Learning Semantic Structure through First-Order-Logic Translation
by: Chaturvedi, Akshay, et al.
Published: (2024)
by: Chaturvedi, Akshay, et al.
Published: (2024)
Realization of Practical Eightfold Fermions and Fourfold van Hove Singularity in TaCo$_2$Te$_2$
by: Rong, Hongtao, et al.
Published: (2022)
by: Rong, Hongtao, et al.
Published: (2022)
Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classification
by: Udagawa, Takuma, et al.
Published: (2025)
by: Udagawa, Takuma, et al.
Published: (2025)
Strong hallucinations from negation and how to fix them
by: Asher, Nicholas, et al.
Published: (2024)
by: Asher, Nicholas, et al.
Published: (2024)
TaCo: Data-adaptive and Query-aware Subspace Collision for High-dimensional Approximate Nearest Neighbor Search
by: Wei, Jiuqi, et al.
Published: (2026)
by: Wei, Jiuqi, et al.
Published: (2026)
Protecting Privacy in Classifiers by Token Manipulation
by: Harel, Re'em, et al.
Published: (2024)
by: Harel, Re'em, et al.
Published: (2024)
CGCE: Classifier-Guided Concept Erasure in Generative Models
by: Nguyen, Viet, et al.
Published: (2025)
by: Nguyen, Viet, et al.
Published: (2025)
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
by: Xue, Yuyang, et al.
Published: (2025)
by: Xue, Yuyang, et al.
Published: (2025)
Nebula: A discourse aware Minecraft Builder
by: Chaturvedi, Akshay, et al.
Published: (2024)
by: Chaturvedi, Akshay, et al.
Published: (2024)
Re-examining learning linear functions in context
by: Naim, Omar, et al.
Published: (2024)
by: Naim, Omar, et al.
Published: (2024)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
On Explaining with Attention Matrices
by: Naim, Omar, et al.
Published: (2024)
by: Naim, Omar, et al.
Published: (2024)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
by: Alqurnawi, Yahia, et al.
Published: (2026)
by: Alqurnawi, Yahia, et al.
Published: (2026)
Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts
by: Amara, Ibtihel, et al.
Published: (2025)
by: Amara, Ibtihel, et al.
Published: (2025)
SSA: Improving Performance With a Better Scoring Function
by: Naim, Omar, et al.
Published: (2025)
by: Naim, Omar, et al.
Published: (2025)
Llamipa: An Incremental Discourse Parser
by: Thompson, Kate, et al.
Published: (2024)
by: Thompson, Kate, et al.
Published: (2024)
Type Theory With Erasure
by: Theocharis, Constantine, et al.
Published: (2026)
by: Theocharis, Constantine, et al.
Published: (2026)
ChatGPT in Linear Algebra: Strides Forward, Steps to Go
by: Bagno, Eli, et al.
Published: (2024)
by: Bagno, Eli, et al.
Published: (2024)
FAME: Fictional Actors for Multilingual Erasure
by: Savelli, Claudio, et al.
Published: (2025)
by: Savelli, Claudio, et al.
Published: (2025)
Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models
by: Fuchi, Masane, et al.
Published: (2025)
by: Fuchi, Masane, et al.
Published: (2025)
Embedded Named Entity Recognition using Probing Classifiers
by: Popovič, Nicholas, et al.
Published: (2024)
by: Popovič, Nicholas, et al.
Published: (2024)
Similar Items
-
ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability
by: Poché, Antonin, et al.
Published: (2025) -
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
by: Jourdan, Fanny
Published: (2024) -
Linear Adversarial Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022) -
TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes
by: Upadhayay, Bibek, et al.
Published: (2023) -
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)