Backdoor for Debias: Mitigating Model Bias with Backdoor Attack-based Artificial Bias
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Shangxi, He, Qiuyang, Yu, Jian, Sang, Jitao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Disguised Wolf Is More Harmful Than a Toothless Tiger: Adaptive Malicious Code Injection Backdoor Attack Leveraging User Behavior as Triggers
von: Wu, Shangxi, et al.
Veröffentlicht: (2024)
von: Wu, Shangxi, et al.
Veröffentlicht: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
Mitigating Gender Bias in Depression Detection via Counterfactual Inference
von: Hu, Mingxuan, et al.
Veröffentlicht: (2025)
von: Hu, Mingxuan, et al.
Veröffentlicht: (2025)
Unmasking Bias in AI: A Systematic Review of Bias Detection and Mitigation Strategies in Electronic Health Record-based Models
von: Chen, Feng, et al.
Veröffentlicht: (2023)
von: Chen, Feng, et al.
Veröffentlicht: (2023)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
Effective Controllable Bias Mitigation for Classification and Retrieval using Gate Adapters
von: Masoudian, Shahed, et al.
Veröffentlicht: (2024)
von: Masoudian, Shahed, et al.
Veröffentlicht: (2024)
Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2024)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2024)
Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
Inference-Time Rule Eraser: Fair Recognition via Distilling and Removing Biased Rules
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
Backdooring Bias ($B^2$) into Stable Diffusion Models
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
Integrating Social Determinants of Health into Knowledge Graphs: Evaluating Prediction Bias and Fairness in Healthcare
von: Shang, Tianqi, et al.
Veröffentlicht: (2024)
von: Shang, Tianqi, et al.
Veröffentlicht: (2024)
Epistemic Uncertainty-Weighted Loss for Visual Bias Mitigation
von: Stone, Rebecca S, et al.
Veröffentlicht: (2022)
von: Stone, Rebecca S, et al.
Veröffentlicht: (2022)
Bias and Fairness in Large Language Models: A Survey
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2023)
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2023)
Say My Name: a Model's Bias Discovery Framework
von: Ciranni, Massimiliano, et al.
Veröffentlicht: (2024)
von: Ciranni, Massimiliano, et al.
Veröffentlicht: (2024)
A Vision-Language Pre-training Model-Guided Approach for Mitigating Backdoor Attacks in Federated Learning
von: Gai, Keke, et al.
Veröffentlicht: (2025)
von: Gai, Keke, et al.
Veröffentlicht: (2025)
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
von: Shvartzshnaider, Yan, et al.
Veröffentlicht: (2024)
von: Shvartzshnaider, Yan, et al.
Veröffentlicht: (2024)
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
von: Cohen-Inger, Nurit, et al.
Veröffentlicht: (2025)
von: Cohen-Inger, Nurit, et al.
Veröffentlicht: (2025)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
von: Abels, Axel, et al.
Veröffentlicht: (2025)
von: Abels, Axel, et al.
Veröffentlicht: (2025)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
von: Salinas, Alejandro, et al.
Veröffentlicht: (2024)
von: Salinas, Alejandro, et al.
Veröffentlicht: (2024)
Breaking Down Bias: On The Limits of Generalizable Pruning Strategies
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions
von: Sánchez-Cortés, Dairazalia, et al.
Veröffentlicht: (2024)
von: Sánchez-Cortés, Dairazalia, et al.
Veröffentlicht: (2024)
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
Impacts of Racial Bias in Historical Training Data for News AI
von: Bhargava, Rahul, et al.
Veröffentlicht: (2025)
von: Bhargava, Rahul, et al.
Veröffentlicht: (2025)
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
von: Porikli, Sedat, et al.
Veröffentlicht: (2025)
von: Porikli, Sedat, et al.
Veröffentlicht: (2025)
Addressing Selection Bias in Computerized Adaptive Testing: A User-Wise Aggregate Influence Function Approach
von: Kwon, Soonwoo, et al.
Veröffentlicht: (2023)
von: Kwon, Soonwoo, et al.
Veröffentlicht: (2023)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
Automatic Assessment of Students' Classroom Engagement with Bias Mitigated Multi-task Model
von: Thiering, James, et al.
Veröffentlicht: (2025)
von: Thiering, James, et al.
Veröffentlicht: (2025)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Demographic Bias of Expert-Level Vision-Language Foundation Models in Medical Imaging
von: Yang, Yuzhe, et al.
Veröffentlicht: (2024)
von: Yang, Yuzhe, et al.
Veröffentlicht: (2024)
Defending Deep Regression Models against Backdoor Attacks
von: Du, Lingyu, et al.
Veröffentlicht: (2024)
von: Du, Lingyu, et al.
Veröffentlicht: (2024)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Disguised Wolf Is More Harmful Than a Toothless Tiger: Adaptive Malicious Code Injection Backdoor Attack Leveraging User Behavior as Triggers
von: Wu, Shangxi, et al.
Veröffentlicht: (2024) -
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023) -
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025) -
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025) -
Mitigating Gender Bias in Depression Detection via Counterfactual Inference
von: Hu, Mingxuan, et al.
Veröffentlicht: (2025)