Backdooring Bias in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Das, Anudeep, Chantasantitam, Prach, Singh, Gurjot, He, Lipeng, Ponomarenko, Mariia, Kerschbaum, Florian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
di: Das, Anudeep, et al.
Pubblicazione: (2025)
di: Das, Anudeep, et al.
Pubblicazione: (2025)
Advanced Real-Time Fraud Detection Using RAG-Based LLMs
di: Singh, Gurjot, et al.
Pubblicazione: (2025)
di: Singh, Gurjot, et al.
Pubblicazione: (2025)
PAL*M: Property Attestation for Large Generative Models
di: Chantasantitam, Prach, et al.
Pubblicazione: (2026)
di: Chantasantitam, Prach, et al.
Pubblicazione: (2026)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
Injecting Bias into Text Classification Models using Backdoor Attacks
di: Yavuz, A. Dilara, et al.
Pubblicazione: (2024)
di: Yavuz, A. Dilara, et al.
Pubblicazione: (2024)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
di: Yu, Miao, et al.
Pubblicazione: (2025)
di: Yu, Miao, et al.
Pubblicazione: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
di: Li, Xi, et al.
Pubblicazione: (2024)
di: Li, Xi, et al.
Pubblicazione: (2024)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
di: Li, Yuetai, et al.
Pubblicazione: (2024)
di: Li, Yuetai, et al.
Pubblicazione: (2024)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
di: Zhou, Yihe, et al.
Pubblicazione: (2025)
di: Zhou, Yihe, et al.
Pubblicazione: (2025)
Concept-Guided Backdoor Attack on Vision Language Models
di: Shen, Haoyu, et al.
Pubblicazione: (2025)
di: Shen, Haoyu, et al.
Pubblicazione: (2025)
Backdooring Bias ($B^2$) into Stable Diffusion Models
di: Naseh, Ali, et al.
Pubblicazione: (2024)
di: Naseh, Ali, et al.
Pubblicazione: (2024)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
di: Truong, Vu Tuan, et al.
Pubblicazione: (2026)
di: Truong, Vu Tuan, et al.
Pubblicazione: (2026)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
di: Cao, Yuanpu, et al.
Pubblicazione: (2023)
di: Cao, Yuanpu, et al.
Pubblicazione: (2023)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
di: Wang, Bingzheng, et al.
Pubblicazione: (2026)
di: Wang, Bingzheng, et al.
Pubblicazione: (2026)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
di: Jiang, Peihai, et al.
Pubblicazione: (2025)
di: Jiang, Peihai, et al.
Pubblicazione: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
Lightweight and Fast Backdoor Model Detection
di: Yu, Yinbo, et al.
Pubblicazione: (2026)
di: Yu, Yinbo, et al.
Pubblicazione: (2026)
Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
di: Peng, Zuquan, et al.
Pubblicazione: (2025)
di: Peng, Zuquan, et al.
Pubblicazione: (2025)
BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
di: Zeng, Yi, et al.
Pubblicazione: (2024)
di: Zeng, Yi, et al.
Pubblicazione: (2024)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
Exploiting the Vulnerability of Large Language Models via Defense-Aware Architectural Backdoor
di: Miah, Abdullah Arafat, et al.
Pubblicazione: (2024)
di: Miah, Abdullah Arafat, et al.
Pubblicazione: (2024)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
di: Wu, Zongru, et al.
Pubblicazione: (2024)
di: Wu, Zongru, et al.
Pubblicazione: (2024)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
di: Yu, Haiyang, et al.
Pubblicazione: (2024)
di: Yu, Haiyang, et al.
Pubblicazione: (2024)
Towards Backdoor Stealthiness in Model Parameter Space
di: Xu, Xiaoyun, et al.
Pubblicazione: (2025)
di: Xu, Xiaoyun, et al.
Pubblicazione: (2025)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
di: Li, Yige, et al.
Pubblicazione: (2026)
di: Li, Yige, et al.
Pubblicazione: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
di: Guo, Weiyang, et al.
Pubblicazione: (2026)
di: Guo, Weiyang, et al.
Pubblicazione: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
di: Li, Yige, et al.
Pubblicazione: (2025)
di: Li, Yige, et al.
Pubblicazione: (2025)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
di: Das, Badhan Chandra, et al.
Pubblicazione: (2026)
di: Das, Badhan Chandra, et al.
Pubblicazione: (2026)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
di: Zheng, Jingyi, et al.
Pubblicazione: (2024)
di: Zheng, Jingyi, et al.
Pubblicazione: (2024)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
di: Riaño, Roberto, et al.
Pubblicazione: (2024)
di: Riaño, Roberto, et al.
Pubblicazione: (2024)
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
di: Braun, Tobias, et al.
Pubblicazione: (2026)
di: Braun, Tobias, et al.
Pubblicazione: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
di: Xu, Zonghuan, et al.
Pubblicazione: (2025)
di: Xu, Zonghuan, et al.
Pubblicazione: (2025)
Stealthy Backdoor Attack to Real-world Models in Android Apps
di: Wei, Jiali, et al.
Pubblicazione: (2025)
di: Wei, Jiali, et al.
Pubblicazione: (2025)
On the Weaknesses of Backdoor-based Model Watermarking: An Information-theoretic Perspective
di: Hu, Aoting, et al.
Pubblicazione: (2024)
di: Hu, Aoting, et al.
Pubblicazione: (2024)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
di: Ma, Yuan, et al.
Pubblicazione: (2024)
di: Ma, Yuan, et al.
Pubblicazione: (2024)
SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation
di: Yang, He, et al.
Pubblicazione: (2026)
di: Yang, He, et al.
Pubblicazione: (2026)
The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models
di: Wu, Zihui, et al.
Pubblicazione: (2024)
di: Wu, Zihui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
di: Das, Anudeep, et al.
Pubblicazione: (2025) -
Advanced Real-Time Fraud Detection Using RAG-Based LLMs
di: Singh, Gurjot, et al.
Pubblicazione: (2025) -
PAL*M: Property Attestation for Large Generative Models
di: Chantasantitam, Prach, et al.
Pubblicazione: (2026) -
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
di: Liu, Xuxu, et al.
Pubblicazione: (2025) -
Injecting Bias into Text Classification Models using Backdoor Attacks
di: Yavuz, A. Dilara, et al.
Pubblicazione: (2024)