BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
Fuente:
arXiv
Saved in:
| Main Authors: | Xue, Jiaqi, Lou, Qian, Zheng, Mengxin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
by: Lou, Qian, et al.
Published: (2024)
by: Lou, Qian, et al.
Published: (2024)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
by: Xiao, Yuxin, et al.
Published: (2023)
by: Xiao, Yuxin, et al.
Published: (2023)
Position Paper: Assessing Robustness, Privacy, and Fairness in Federated Learning Integrated with Foundation Models
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
FairDP: Certified Fairness with Differential Privacy
by: Tran, Khang, et al.
Published: (2023)
by: Tran, Khang, et al.
Published: (2023)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
RobPI: Robust Private Inference against Malicious Client
by: Xue, Jiaqi, et al.
Published: (2026)
by: Xue, Jiaqi, et al.
Published: (2026)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
Differentially Private Post-Processing for Fair Regression
by: Xian, Ruicheng, et al.
Published: (2024)
by: Xian, Ruicheng, et al.
Published: (2024)
Learning Fair Robustness via Domain Mixup
by: Zhong, Meiyu, et al.
Published: (2024)
by: Zhong, Meiyu, et al.
Published: (2024)
Privacy Constrained Fairness Estimation for Decision Trees
by: van der Steen, Florian, et al.
Published: (2023)
by: van der Steen, Florian, et al.
Published: (2023)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
On the Impact of Multi-dimensional Local Differential Privacy on Fairness
by: Makhlouf, Karima, et al.
Published: (2023)
by: Makhlouf, Karima, et al.
Published: (2023)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
by: Yao, Duanyi, et al.
Published: (2026)
by: Yao, Duanyi, et al.
Published: (2026)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
TFHE-Coder: Evaluating LLM-agentic Fully Homomorphic Encryption Code Generation
by: Kumar, Mayank, et al.
Published: (2025)
by: Kumar, Mayank, et al.
Published: (2025)
Backdoor Attack with Sparse and Invisible Trigger
by: Gao, Yinghua, et al.
Published: (2023)
by: Gao, Yinghua, et al.
Published: (2023)
RPP: A Certified Poisoned-Sample Detection Framework for Backdoor Attacks under Dataset Imbalance
by: Lin, Miao, et al.
Published: (2026)
by: Lin, Miao, et al.
Published: (2026)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
by: He, Jiaming, et al.
Published: (2024)
by: He, Jiaming, et al.
Published: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
by: Liu, Qin, et al.
Published: (2023)
by: Liu, Qin, et al.
Published: (2023)
PUFFLE: Balancing Privacy, Utility, and Fairness in Federated Learning
by: Corbucci, Luca, et al.
Published: (2024)
by: Corbucci, Luca, et al.
Published: (2024)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
by: Jang, Yeonwoo, et al.
Published: (2025)
by: Jang, Yeonwoo, et al.
Published: (2025)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024)
by: Tong, Terry, et al.
Published: (2024)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Privacy at a Price: Exploring its Dual Impact on AI Fairness
by: Yang, Mengmeng, et al.
Published: (2024)
by: Yang, Mengmeng, et al.
Published: (2024)
FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning
by: Sanchez Jr., David, et al.
Published: (2025)
by: Sanchez Jr., David, et al.
Published: (2025)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
by: Hou, Abe Bohan, et al.
Published: (2024)
by: Hou, Abe Bohan, et al.
Published: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
by: Zhang, Zheng, et al.
Published: (2024)
by: Zhang, Zheng, et al.
Published: (2024)
SSL-Cleanse: Trojan Detection and Mitigation in Self-Supervised Learning
by: Zheng, Mengxin, et al.
Published: (2023)
by: Zheng, Mengxin, et al.
Published: (2023)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
by: Yan, Jun, et al.
Published: (2023)
by: Yan, Jun, et al.
Published: (2023)
Test-Time Backdoor Attacks on Multimodal Large Language Models
by: Lu, Dong, et al.
Published: (2024)
by: Lu, Dong, et al.
Published: (2024)
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024)
by: Zhang, Jinghuai, et al.
Published: (2024)
Similar Items
-
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
by: Xue, Jiaqi, et al.
Published: (2024) -
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024) -
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
by: Lou, Qian, et al.
Published: (2024) -
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
by: Xiao, Yuxin, et al.
Published: (2023) -
Position Paper: Assessing Robustness, Privacy, and Fairness in Federated Learning Integrated with Foundation Models
by: Wang, Jiaqi, et al.
Published: (2024)