Injecting Bias into Text Classification Models using Backdoor Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Yavuz, A. Dilara, Gursoy, M. Emre |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
by: Kirci, Onur Alp, et al.
Published: (2025)
by: Kirci, Onur Alp, et al.
Published: (2025)
Defending Against Beta Poisoning Attacks in Machine Learning Models
by: Gulciftci, Nilufer, et al.
Published: (2025)
by: Gulciftci, Nilufer, et al.
Published: (2025)
Win-k: Improved Membership Inference Attacks on Small Language Models
by: Arkhmammadova, Roya, et al.
Published: (2025)
by: Arkhmammadova, Roya, et al.
Published: (2025)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
by: Chen, Tianxin, et al.
Published: (2026)
by: Chen, Tianxin, et al.
Published: (2026)
Backdooring Bias in Large Language Models
by: Das, Anudeep, et al.
Published: (2026)
by: Das, Anudeep, et al.
Published: (2026)
Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs
by: Pallakonda, Bhanu, et al.
Published: (2026)
by: Pallakonda, Bhanu, et al.
Published: (2026)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
by: An, Shengwei, et al.
Published: (2023)
by: An, Shengwei, et al.
Published: (2023)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
Concept-Guided Backdoor Attack on Vision Language Models
by: Shen, Haoyu, et al.
Published: (2025)
by: Shen, Haoyu, et al.
Published: (2025)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
by: Popovic, Dorde, et al.
Published: (2025)
by: Popovic, Dorde, et al.
Published: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
by: Li, Xi, et al.
Published: (2024)
by: Li, Xi, et al.
Published: (2024)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
by: Ma, Yuan, et al.
Published: (2024)
by: Ma, Yuan, et al.
Published: (2024)
Stealthy Backdoor Attack to Real-world Models in Android Apps
by: Wei, Jiali, et al.
Published: (2025)
by: Wei, Jiali, et al.
Published: (2025)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
by: Riaño, Roberto, et al.
Published: (2024)
by: Riaño, Roberto, et al.
Published: (2024)
VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models
by: Liu, Pang, et al.
Published: (2026)
by: Liu, Pang, et al.
Published: (2026)
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
by: Zou, Wei, et al.
Published: (2026)
by: Zou, Wei, et al.
Published: (2026)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
by: Song, Baogang, et al.
Published: (2025)
by: Song, Baogang, et al.
Published: (2025)
Revisiting Backdoor Attacks on Time Series Classification in the Frequency Domain
by: Huang, Yuanmin, et al.
Published: (2025)
by: Huang, Yuanmin, et al.
Published: (2025)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
by: Ding, Zikang, et al.
Published: (2026)
by: Ding, Zikang, et al.
Published: (2026)
Optimal Defenses Against Gradient Reconstruction Attacks
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
by: Li, Yuetai, et al.
Published: (2024)
by: Li, Yuetai, et al.
Published: (2024)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
by: Tie, Guiyao, et al.
Published: (2026)
by: Tie, Guiyao, et al.
Published: (2026)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
by: Liu, Xuxu, et al.
Published: (2025)
by: Liu, Xuxu, et al.
Published: (2025)
Towards Effective, Stealthy, and Persistent Backdoor Attacks Targeting Graph Foundation Models
by: Luo, Jiayi, et al.
Published: (2025)
by: Luo, Jiayi, et al.
Published: (2025)
Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models
by: Chen, Linzhi, et al.
Published: (2025)
by: Chen, Linzhi, et al.
Published: (2025)
Clean-Label Physical Backdoor Attacks with Data Distillation
by: Dao, Thinh, et al.
Published: (2024)
by: Dao, Thinh, et al.
Published: (2024)
Invisible Textual Backdoor Attacks based on Dual-Trigger
by: Hou, Yang, et al.
Published: (2024)
by: Hou, Yang, et al.
Published: (2024)
Impart: An Imperceptible and Effective Label-Specific Backdoor Attack
by: Zhao, Jingke, et al.
Published: (2024)
by: Zhao, Jingke, et al.
Published: (2024)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025)
by: Lin, Chenhao, et al.
Published: (2025)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
by: Ning, Liangbo, et al.
Published: (2025)
by: Ning, Liangbo, et al.
Published: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
by: Qiu, Huming, et al.
Published: (2023)
by: Qiu, Huming, et al.
Published: (2023)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
by: Zhou, Yihe, et al.
Published: (2025)
by: Zhou, Yihe, et al.
Published: (2025)
Context is the Key: Backdoor Attacks for In-Context Learning with Vision Transformers
by: Abad, Gorka, et al.
Published: (2024)
by: Abad, Gorka, et al.
Published: (2024)
CBPF: Filtering Poisoned Data Based on Composite Backdoor Attack
by: Xia, Hanfeng, et al.
Published: (2024)
by: Xia, Hanfeng, et al.
Published: (2024)
Similar Items
-
Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
by: Kirci, Onur Alp, et al.
Published: (2025) -
Defending Against Beta Poisoning Attacks in Machine Learning Models
by: Gulciftci, Nilufer, et al.
Published: (2025) -
Win-k: Improved Membership Inference Attacks on Small Language Models
by: Arkhmammadova, Roya, et al.
Published: (2025) -
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
by: Chen, Tianxin, et al.
Published: (2026) -
Backdooring Bias in Large Language Models
by: Das, Anudeep, et al.
Published: (2026)