Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
Fuente:
arXiv
Salvato in:
| Autori principali: | Guo, Weiyang, Shi, Zesheng, Zhu, Zeen, Zhou, Yuan, Zhang, Min, Li, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
di: Shen, Guangyu, et al.
Pubblicazione: (2025)
di: Shen, Guangyu, et al.
Pubblicazione: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
di: Chen, Zhuowei, et al.
Pubblicazione: (2025)
di: Chen, Zhuowei, et al.
Pubblicazione: (2025)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
di: Li, Yige, et al.
Pubblicazione: (2026)
di: Li, Yige, et al.
Pubblicazione: (2026)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
di: Ji, Wence, et al.
Pubblicazione: (2025)
di: Ji, Wence, et al.
Pubblicazione: (2025)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
di: Li, Yige, et al.
Pubblicazione: (2025)
di: Li, Yige, et al.
Pubblicazione: (2025)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
di: Cui, Jing, et al.
Pubblicazione: (2025)
di: Cui, Jing, et al.
Pubblicazione: (2025)
Lightweight and Fast Backdoor Model Detection
di: Yu, Yinbo, et al.
Pubblicazione: (2026)
di: Yu, Yinbo, et al.
Pubblicazione: (2026)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
di: Yu, Miao, et al.
Pubblicazione: (2025)
di: Yu, Miao, et al.
Pubblicazione: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
di: Qiu, Huming, et al.
Pubblicazione: (2023)
di: Qiu, Huming, et al.
Pubblicazione: (2023)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
di: Wang, Bingzheng, et al.
Pubblicazione: (2026)
di: Wang, Bingzheng, et al.
Pubblicazione: (2026)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
di: Yu, Haiyang, et al.
Pubblicazione: (2024)
di: Yu, Haiyang, et al.
Pubblicazione: (2024)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
di: Zhang, Rui, et al.
Pubblicazione: (2025)
di: Zhang, Rui, et al.
Pubblicazione: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
di: Jiang, Peihai, et al.
Pubblicazione: (2025)
di: Jiang, Peihai, et al.
Pubblicazione: (2025)
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
di: Wan, Wei, et al.
Pubblicazione: (2025)
di: Wan, Wei, et al.
Pubblicazione: (2025)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
di: Hu, Man, et al.
Pubblicazione: (2025)
di: Hu, Man, et al.
Pubblicazione: (2025)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
di: Zhou, Yihe, et al.
Pubblicazione: (2025)
di: Zhou, Yihe, et al.
Pubblicazione: (2025)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
di: Riaño, Roberto, et al.
Pubblicazione: (2024)
di: Riaño, Roberto, et al.
Pubblicazione: (2024)
Repurposing Backdoors for Good: Ephemeral Intrinsic Proofs for Verifiable Aggregation in Cross-silo Federated Learning
di: Qin, Xian, et al.
Pubblicazione: (2026)
di: Qin, Xian, et al.
Pubblicazione: (2026)
WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks
di: Li, Tingzhi, et al.
Pubblicazione: (2025)
di: Li, Tingzhi, et al.
Pubblicazione: (2025)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
di: Tie, Guiyao, et al.
Pubblicazione: (2026)
di: Tie, Guiyao, et al.
Pubblicazione: (2026)
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
di: Zhou, Qi, et al.
Pubblicazione: (2024)
di: Zhou, Qi, et al.
Pubblicazione: (2024)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
di: Guo, Weiyang, et al.
Pubblicazione: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
di: Wang, Yifei, et al.
Pubblicazione: (2026)
di: Wang, Yifei, et al.
Pubblicazione: (2026)
Universal Jailbreak Backdoors from Poisoned Human Feedback
di: Rando, Javier, et al.
Pubblicazione: (2023)
di: Rando, Javier, et al.
Pubblicazione: (2023)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
Compromising Embodied Agents with Contextual Backdoor Attacks
di: Liu, Aishan, et al.
Pubblicazione: (2024)
di: Liu, Aishan, et al.
Pubblicazione: (2024)
Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
di: Xue, Xiaoyu, et al.
Pubblicazione: (2025)
di: Xue, Xiaoyu, et al.
Pubblicazione: (2025)
DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent
di: Zhu, Pengyu, et al.
Pubblicazione: (2025)
di: Zhu, Pengyu, et al.
Pubblicazione: (2025)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs
di: Pallakonda, Bhanu, et al.
Pubblicazione: (2026)
di: Pallakonda, Bhanu, et al.
Pubblicazione: (2026)
BadEdit: Backdooring large language models by model editing
di: Li, Yanzhou, et al.
Pubblicazione: (2024)
di: Li, Yanzhou, et al.
Pubblicazione: (2024)
Towards Effective, Stealthy, and Persistent Backdoor Attacks Targeting Graph Foundation Models
di: Luo, Jiayi, et al.
Pubblicazione: (2025)
di: Luo, Jiayi, et al.
Pubblicazione: (2025)
You Can Backdoor Personalized Federated Learning
di: Ye, Tiandi, et al.
Pubblicazione: (2023)
di: Ye, Tiandi, et al.
Pubblicazione: (2023)
Dark Distillation: Backdooring Distilled Datasets without Accessing Raw Data
di: Yang, Ziyuan, et al.
Pubblicazione: (2025)
di: Yang, Ziyuan, et al.
Pubblicazione: (2025)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
di: Ma, Yuan, et al.
Pubblicazione: (2024)
di: Ma, Yuan, et al.
Pubblicazione: (2024)
Backdooring Bias in Large Language Models
di: Das, Anudeep, et al.
Pubblicazione: (2026)
di: Das, Anudeep, et al.
Pubblicazione: (2026)
Does Few-shot Learning Suffer from Backdoor Attacks?
di: Liu, Xinwei, et al.
Pubblicazione: (2023)
di: Liu, Xinwei, et al.
Pubblicazione: (2023)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
di: Min, Rui, et al.
Pubblicazione: (2024)
di: Min, Rui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
di: Shen, Guangyu, et al.
Pubblicazione: (2025) -
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
di: Chen, Zhuowei, et al.
Pubblicazione: (2025) -
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
di: Li, Yige, et al.
Pubblicazione: (2026) -
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
di: Ji, Wence, et al.
Pubblicazione: (2025) -
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
di: Li, Yige, et al.
Pubblicazione: (2025)