Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jianan, Qin, Simeng, Jia, Xiaojun, Wang, Lionel Z., Zheng, Tianhang, Jia, Xiaoshuang, Liu, Yang, Cao, Xiaochun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
by: Cheng, Ruoxi, et al.
Published: (2024)
by: Cheng, Ruoxi, et al.
Published: (2024)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
ODAR: Principled Adaptive Routing for LLM Reasoning via Active Inference
by: Ma, Siyuan, et al.
Published: (2026)
by: Ma, Siyuan, et al.
Published: (2026)
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
by: Huang, Xun, et al.
Published: (2026)
by: Huang, Xun, et al.
Published: (2026)
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
by: Liu, Xinwei, et al.
Published: (2025)
by: Liu, Xinwei, et al.
Published: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
by: Cheng, Ruoxi, et al.
Published: (2025)
by: Cheng, Ruoxi, et al.
Published: (2025)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
by: Teng, Ma, et al.
Published: (2024)
by: Teng, Ma, et al.
Published: (2024)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
by: Jia, Xiaojun, et al.
Published: (2026)
by: Jia, Xiaojun, et al.
Published: (2026)
Does Few-shot Learning Suffer from Backdoor Attacks?
by: Liu, Xinwei, et al.
Published: (2023)
by: Liu, Xinwei, et al.
Published: (2023)
PersGuard: Preventing Malicious Personalization via Backdoor Attacks on Pre-trained Text-to-Image Diffusion Models
by: Liu, Xinwei, et al.
Published: (2025)
by: Liu, Xinwei, et al.
Published: (2025)
Dynamic Target Attack
by: Xiu, Kedong, et al.
Published: (2025)
by: Xiu, Kedong, et al.
Published: (2025)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
by: Yang, Cehao, et al.
Published: (2025)
by: Yang, Cehao, et al.
Published: (2025)
SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning
by: Yao, Jian, et al.
Published: (2026)
by: Yao, Jian, et al.
Published: (2026)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
by: Lan, Yifan, et al.
Published: (2026)
by: Lan, Yifan, et al.
Published: (2026)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
by: Liang, Jia, et al.
Published: (2026)
by: Liang, Jia, et al.
Published: (2026)
Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
by: Lyu, Tianwen, et al.
Published: (2025)
by: Lyu, Tianwen, et al.
Published: (2025)
Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning
by: Hong, Jialiang, et al.
Published: (2025)
by: Hong, Jialiang, et al.
Published: (2025)
Evaluating LLM Reasoning Beyond Correctness and CoT
by: Abbasloo, Soheil
Published: (2025)
by: Abbasloo, Soheil
Published: (2025)
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
by: Chen, Guangke, et al.
Published: (2025)
by: Chen, Guangke, et al.
Published: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Stepwise Reasoning Error Disruption Attack of LLMs
by: Peng, Jingyu, et al.
Published: (2024)
by: Peng, Jingyu, et al.
Published: (2024)
CoT2-Meta: Budgeted Metacognitive Control for Test-Time Reasoning
by: Ma, Siyuan, et al.
Published: (2026)
by: Ma, Siyuan, et al.
Published: (2026)
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
by: Cao, Chentao, et al.
Published: (2025)
by: Cao, Chentao, et al.
Published: (2025)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
by: Xun, Yuan, et al.
Published: (2024)
by: Xun, Yuan, et al.
Published: (2024)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
by: Ren, Ruifeng, et al.
Published: (2024)
by: Ren, Ruifeng, et al.
Published: (2024)
RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
by: Chen, Jianhao, et al.
Published: (2025)
by: Chen, Jianhao, et al.
Published: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Efficient Long CoT Reasoning in Small Language Models
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics
by: Bachmann, Gregor, et al.
Published: (2026)
by: Bachmann, Gregor, et al.
Published: (2026)
Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
by: Luo, Yijia, et al.
Published: (2025)
by: Luo, Yijia, et al.
Published: (2025)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
by: Zhou, Weibo, et al.
Published: (2025)
by: Zhou, Weibo, et al.
Published: (2025)
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
by: Kumarage, Tharindu, et al.
Published: (2025)
by: Kumarage, Tharindu, et al.
Published: (2025)
Similar Items
-
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
by: Cheng, Ruoxi, et al.
Published: (2024) -
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025) -
ODAR: Principled Adaptive Routing for LLM Reasoning via Active Inference
by: Ma, Siyuan, et al.
Published: (2026) -
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
by: Huang, Xun, et al.
Published: (2026) -
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
by: Liu, Xinwei, et al.
Published: (2025)