Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Man, Wu, Xinyi, Suo, Zuofeng, Feng, Jinbo, Meng, Linghui, Jia, Yanhao, Luu, Anh Tuan, Zhao, Shuai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
por: Zhao, Shuai, et al.
Publicado: (2025)
por: Zhao, Shuai, et al.
Publicado: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
por: Hu, Man, et al.
Publicado: (2025)
por: Hu, Man, et al.
Publicado: (2025)
Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models
por: Wen, Jinming, et al.
Publicado: (2025)
por: Wen, Jinming, et al.
Publicado: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
Excessive Reasoning Attack on Reasoning LLMs
por: Si, Wai Man, et al.
Publicado: (2025)
por: Si, Wai Man, et al.
Publicado: (2025)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
por: Truong, Vu Tuan, et al.
Publicado: (2026)
por: Truong, Vu Tuan, et al.
Publicado: (2026)
Rethinking Backdoor Detection Evaluation for Language Models
por: Yan, Jun, et al.
Publicado: (2024)
por: Yan, Jun, et al.
Publicado: (2024)
ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
por: Yan, Lu, et al.
Publicado: (2024)
por: Yan, Lu, et al.
Publicado: (2024)
The Invitation Trap: Proactive Availability Backdoor in LLMs via Conversational Induction
por: Wang, He, et al.
Publicado: (2026)
por: Wang, He, et al.
Publicado: (2026)
TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
por: Guo, Zhen, et al.
Publicado: (2026)
por: Guo, Zhen, et al.
Publicado: (2026)
MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning
por: Zeng, Yizhe, et al.
Publicado: (2026)
por: Zeng, Yizhe, et al.
Publicado: (2026)
When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs
por: Li, Yue, et al.
Publicado: (2025)
por: Li, Yue, et al.
Publicado: (2025)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
por: Li, Yige, et al.
Publicado: (2026)
por: Li, Yige, et al.
Publicado: (2026)
Synthesizing Physical Backdoor Datasets: An Automated Framework Leveraging Deep Generative Models
por: Yang, Sze Jue, et al.
Publicado: (2023)
por: Yang, Sze Jue, et al.
Publicado: (2023)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
por: Zhou, Yihe, et al.
Publicado: (2025)
por: Zhou, Yihe, et al.
Publicado: (2025)
Execution-State-Aware LLM Reasoning for Automated Proof-of-Vulnerability Generation
por: Li, Haoyu, et al.
Publicado: (2026)
por: Li, Haoyu, et al.
Publicado: (2026)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
por: Chen, Chen, et al.
Publicado: (2026)
por: Chen, Chen, et al.
Publicado: (2026)
Adaptive Backdoor Attacks with Reasonable Constraints on Graph Neural Networks
por: Dong, Xuewen, et al.
Publicado: (2025)
por: Dong, Xuewen, et al.
Publicado: (2025)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
por: Wang, Qingyue, et al.
Publicado: (2025)
por: Wang, Qingyue, et al.
Publicado: (2025)
ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data
por: Zhao, Haodong, et al.
Publicado: (2026)
por: Zhao, Haodong, et al.
Publicado: (2026)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
por: Zhu, Zhenhao, et al.
Publicado: (2026)
por: Zhu, Zhenhao, et al.
Publicado: (2026)
Backdoor Attacks and Defenses in Computer Vision Domain: A Survey
por: Abbasi, Bilal Hussain, et al.
Publicado: (2025)
por: Abbasi, Bilal Hussain, et al.
Publicado: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
por: Guo, Weiyang, et al.
Publicado: (2026)
por: Guo, Weiyang, et al.
Publicado: (2026)
Revisiting Backdoor Threat in Federated Instruction Tuning from a Signal Aggregation Perspective
por: Zhao, Haodong, et al.
Publicado: (2026)
por: Zhao, Haodong, et al.
Publicado: (2026)
Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
RHINO: Guided Reasoning for Mapping Network Logs to Adversarial Tactics and Techniques with Large Language Models
por: Meng, Fanchao, et al.
Publicado: (2025)
por: Meng, Fanchao, et al.
Publicado: (2025)
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
por: Fu, Hang, et al.
Publicado: (2026)
por: Fu, Hang, et al.
Publicado: (2026)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
por: Liu, Mingrui, et al.
Publicado: (2025)
por: Liu, Mingrui, et al.
Publicado: (2025)
E-SAGE: Explainability-based Defense Against Backdoor Attacks on Graph Neural Networks
por: Yuan, Dingqiang, et al.
Publicado: (2024)
por: Yuan, Dingqiang, et al.
Publicado: (2024)
Cryptographic Backdoor for Neural Networks: Boon and Bane
por: Ngo, Anh Tu, et al.
Publicado: (2025)
por: Ngo, Anh Tu, et al.
Publicado: (2025)
Architectural Backdoors in Deep Learning: A Survey of Vulnerabilities, Detection, and Defense
por: Childress, Victoria, et al.
Publicado: (2025)
por: Childress, Victoria, et al.
Publicado: (2025)
Stateful Agent Backdoor
por: Dai, Zhengchunmin, et al.
Publicado: (2026)
por: Dai, Zhengchunmin, et al.
Publicado: (2026)
When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
por: Hu, Ruihan, et al.
Publicado: (2026)
por: Hu, Ruihan, et al.
Publicado: (2026)
Activation Gradient based Poisoned Sample Detection Against Backdoor Attacks
por: Yuan, Danni, et al.
Publicado: (2023)
por: Yuan, Danni, et al.
Publicado: (2023)
Rethinking Graph Backdoor Attacks: A Distribution-Preserving Perspective
por: Zhang, Zhiwei, et al.
Publicado: (2024)
por: Zhang, Zhiwei, et al.
Publicado: (2024)
Ejemplares similares
-
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
por: Zhao, Shuai, et al.
Publicado: (2025) -
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
por: Zhao, Shuai, et al.
Publicado: (2024) -
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
por: Zhao, Shuai, et al.
Publicado: (2024) -
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
por: Zhao, Shuai, et al.
Publicado: (2024) -
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
por: Zhao, Shuai, et al.
Publicado: (2024)