SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Fengqing, Xu, Zhangchen, Li, Yuetai, Niu, Luyao, Xiang, Zhen, Li, Bo, Lin, Bill Yuchen, Poovendran, Radha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Temporal Sampling for Forgotten Reasoning in LLMs
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Small Models Struggle to Learn from Strong Reasoners
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Stronger Models are NOT Stronger Teachers for Instruction Tuning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
by: Feng, Yichen, et al.
Published: (2025)
by: Feng, Yichen, et al.
Published: (2025)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
by: Li, Yuetai, et al.
Published: (2024)
by: Li, Yuetai, et al.
Published: (2024)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
by: Xiang, Zhen, et al.
Published: (2024)
by: Xiang, Zhen, et al.
Published: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks
by: Jiang, Fengqing, et al.
Published: (2026)
by: Jiang, Fengqing, et al.
Published: (2026)
Polyhedral Instability Governs Regret in Online Learning
by: Li, Yuetai, et al.
Published: (2026)
by: Li, Yuetai, et al.
Published: (2026)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
Distributed Safety-Critical Control of Multi-Agent Systems with Time-Varying Communication Topologies
by: Cheng, Shiyu, et al.
Published: (2026)
by: Cheng, Shiyu, et al.
Published: (2026)
Demystifying Long Chain-of-Thought Reasoning in LLMs
by: Yeo, Edward, et al.
Published: (2025)
by: Yeo, Edward, et al.
Published: (2025)
Fault Tolerant Neural Control Barrier Functions for Robotic Systems under Sensor Faults and Attacks
by: Zhang, Hongchao, et al.
Published: (2024)
by: Zhang, Hongchao, et al.
Published: (2024)
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
by: Feng, Yichen, et al.
Published: (2026)
by: Feng, Yichen, et al.
Published: (2026)
A Compositional Resilience Index for Computationally Efficient Safety Analysis of Interconnected Systems
by: Niu, Luyao, et al.
Published: (2023)
by: Niu, Luyao, et al.
Published: (2023)
Who is Responsible? Explaining Safety Violations in Multi-Agent Cyber-Physical Systems
by: Niu, Luyao, et al.
Published: (2024)
by: Niu, Luyao, et al.
Published: (2024)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
by: Sahabandu, Dinuka, et al.
Published: (2024)
by: Sahabandu, Dinuka, et al.
Published: (2024)
Long Chain-of-Thought Reasoning Across Languages
by: Barua, Josh, et al.
Published: (2025)
by: Barua, Josh, et al.
Published: (2025)
Modeling and Designing Non-Pharmaceutical Interventions in Epidemics: A Submodular Approach
by: Cheng, Shiyu, et al.
Published: (2024)
by: Cheng, Shiyu, et al.
Published: (2024)
Swarm-STL: A Framework for Motion Planning in Large-Scale, Multi-Swarm Systems
by: Cheng, Shiyu, et al.
Published: (2025)
by: Cheng, Shiyu, et al.
Published: (2025)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
by: Ye, Jiacheng, et al.
Published: (2024)
by: Ye, Jiacheng, et al.
Published: (2024)
Simulating Environments with Reasoning Models for Agent Training
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
by: Lin, Bill Yuchen, et al.
Published: (2025)
by: Lin, Bill Yuchen, et al.
Published: (2025)
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
by: Jiang, Gangwei, et al.
Published: (2025)
by: Jiang, Gangwei, et al.
Published: (2025)
Unveiling Confirmation Bias in Chain-of-Thought Reasoning
by: Wan, Yue, et al.
Published: (2025)
by: Wan, Yue, et al.
Published: (2025)
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
by: He, Yancheng, et al.
Published: (2025)
by: He, Yancheng, et al.
Published: (2025)
Fractured Chain-of-Thought Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation Engineering
by: Tang, Xinyu, et al.
Published: (2025)
by: Tang, Xinyu, et al.
Published: (2025)
Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs
by: Cao, Jie, et al.
Published: (2026)
by: Cao, Jie, et al.
Published: (2026)
Similar Items
-
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
by: Xu, Zhangchen, et al.
Published: (2025) -
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024) -
Temporal Sampling for Forgotten Reasoning in LLMs
by: Li, Yuetai, et al.
Published: (2025) -
Small Models Struggle to Learn from Strong Reasoners
by: Li, Yuetai, et al.
Published: (2025) -
Stronger Models are NOT Stronger Teachers for Instruction Tuning
by: Xu, Zhangchen, et al.
Published: (2024)