H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kuo, Martin, Zhang, Jianyi, Ding, Aolin, Wang, Qinsi, DiValentin, Louis, Bao, Yujia, Wei, Wei, Li, Hai, Chen, Yiran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
von: Hu, Yinghao, et al.
Veröffentlicht: (2025)
von: Hu, Yinghao, et al.
Veröffentlicht: (2025)
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model Utility
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
von: Sands, Brendan, et al.
Veröffentlicht: (2025)
von: Sands, Brendan, et al.
Veröffentlicht: (2025)
Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
von: Lu, Chengda, et al.
Veröffentlicht: (2025)
von: Lu, Chengda, et al.
Veröffentlicht: (2025)
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
von: Mondal, Debjyoti, et al.
Veröffentlicht: (2024)
von: Mondal, Debjyoti, et al.
Veröffentlicht: (2024)
CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning
von: Fang, Yuanheng, et al.
Veröffentlicht: (2025)
von: Fang, Yuanheng, et al.
Veröffentlicht: (2025)
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
Chain-of-Thought Hijacking
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
von: Moell, Birger, et al.
Veröffentlicht: (2025)
von: Moell, Birger, et al.
Veröffentlicht: (2025)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
von: Chua, James, et al.
Veröffentlicht: (2025)
von: Chua, James, et al.
Veröffentlicht: (2025)
Comparative Analysis of OpenAI GPT-4o and DeepSeek R1 for Scientific Text Categorization Using Prompt Engineering
von: Maiti, Aniruddha, et al.
Veröffentlicht: (2025)
von: Maiti, Aniruddha, et al.
Veröffentlicht: (2025)
FedProphet: Memory-Efficient Federated Adversarial Training via Robust and Consistent Cascade Learning
von: Tang, Minxue, et al.
Veröffentlicht: (2024)
von: Tang, Minxue, et al.
Veröffentlicht: (2024)
Accuracy of ChatGPT , Gemini, Claude and DeepSeek in Carbohydrate Counting
von: Luca Zagaroli, et al.
Veröffentlicht: (2026)
von: Luca Zagaroli, et al.
Veröffentlicht: (2026)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
von: Shmidman, Shaltiel, et al.
Veröffentlicht: (2025)
von: Shmidman, Shaltiel, et al.
Veröffentlicht: (2025)
Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor
von: Chang, Wenhan, et al.
Veröffentlicht: (2026)
von: Chang, Wenhan, et al.
Veröffentlicht: (2026)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Co-CoT: A Prompt-Based Framework for Collaborative Chain-of-Thought Reasoning
von: Yoo, Seunghyun
Veröffentlicht: (2025)
von: Yoo, Seunghyun
Veröffentlicht: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning
von: Chen, Shuxu, et al.
Veröffentlicht: (2026)
von: Chen, Shuxu, et al.
Veröffentlicht: (2026)
Chain-of-Sanitized-Thoughts: Plugging PII Leakage in CoT of Large Reasoning Models
von: Das, Arghyadeep, et al.
Veröffentlicht: (2026)
von: Das, Arghyadeep, et al.
Veröffentlicht: (2026)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
Considerations on GPT ‐4o and Gemini Flash 2.0 in Acne/Rosacea
von: Lien‐Chung Wei, et al.
Veröffentlicht: (2025)
von: Lien‐Chung Wei, et al.
Veröffentlicht: (2025)
Who Knows Anatomy Best? A Comparative Study of ChatGPT ‐4o, DeepSeek , Gemini, and Claude
von: Melek Tassoker
Veröffentlicht: (2025)
von: Melek Tassoker
Veröffentlicht: (2025)
SIM-CoT: Supervised Implicit Chain-of-Thought
von: Wei, Xilin, et al.
Veröffentlicht: (2025)
von: Wei, Xilin, et al.
Veröffentlicht: (2025)
V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
von: Xu, Pusheng, et al.
Veröffentlicht: (2025) -
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
von: Hu, Yinghao, et al.
Veröffentlicht: (2025) -
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues
von: Kuo, Martin, et al.
Veröffentlicht: (2025) -
Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model Utility
von: Kuo, Martin, et al.
Veröffentlicht: (2025) -
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
von: Sands, Brendan, et al.
Veröffentlicht: (2025)