The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Yonghong, Yang, Zhen, Jian, Ping, Zhang, Xinyue, Guo, Zhongbin, Li, Chengzhi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
by: Yang, Zhen, et al.
Published: (2026)
by: Yang, Zhen, et al.
Published: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025)
by: Du, Hongzhe, et al.
Published: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
by: Zhou, Huilin, et al.
Published: (2026)
by: Zhou, Huilin, et al.
Published: (2026)
A Closer Look at Adversarial Suffix Learning for Jailbreaking LLMs: Augmented Adversarial Trigger Learning
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
by: Guo, Zhongbin, et al.
Published: (2026)
by: Guo, Zhongbin, et al.
Published: (2026)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
by: Wu, Tongxi, et al.
Published: (2026)
by: Wu, Tongxi, et al.
Published: (2026)
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
by: Peng, Ze, et al.
Published: (2025)
by: Peng, Ze, et al.
Published: (2025)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
by: Cheng, Stephen, et al.
Published: (2026)
by: Cheng, Stephen, et al.
Published: (2026)
Jailbreaking LLMs via Calibration
by: Lu, Yuxuan, et al.
Published: (2026)
by: Lu, Yuxuan, et al.
Published: (2026)
Federated Continual Instruction Tuning
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
Neural Predictive Control to Coordinate Discrete- and Continuous-Time Models for Time-Series Analysis with Control-Theoretical Improvements
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
A Systematic Investigation of The RL-Jailbreaker in LLMs
by: Mohammedalamen, Montaser, et al.
Published: (2026)
by: Mohammedalamen, Montaser, et al.
Published: (2026)
Continued AI Scaling Requires Repeated Efficiency Doublings
by: Lu, Chien-Ping
Published: (2026)
by: Lu, Chien-Ping
Published: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?
by: Karim, Aabid, et al.
Published: (2025)
by: Karim, Aabid, et al.
Published: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
by: Chen, Wenyu, et al.
Published: (2026)
by: Chen, Wenyu, et al.
Published: (2026)
Continual Learning for Smart City: A Survey
by: Yang, Li, et al.
Published: (2024)
by: Yang, Li, et al.
Published: (2024)
Self-Evolving LLMs via Continual Instruction Tuning
by: Kang, Jiazheng, et al.
Published: (2025)
by: Kang, Jiazheng, et al.
Published: (2025)
CoEvo: Continual Evolution of Symbolic Solutions Using Large Language Models
by: Guo, Ping, et al.
Published: (2024)
by: Guo, Ping, et al.
Published: (2024)
Sparse Adapter Fusion for Continual Learning in NLP
by: Zeng, Min, et al.
Published: (2026)
by: Zeng, Min, et al.
Published: (2026)
Information-Theoretic Dual Memory System for Continual Learning
by: Wu, RunQing, et al.
Published: (2025)
by: Wu, RunQing, et al.
Published: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025)
by: Malek, Alan, et al.
Published: (2025)
The Struggles of LLMs in Cross-lingual Code Clone Detection
by: Moumoula, Micheline Bénédicte, et al.
Published: (2024)
by: Moumoula, Micheline Bénédicte, et al.
Published: (2024)
Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMs
by: Zhang, Dingkun, et al.
Published: (2025)
by: Zhang, Dingkun, et al.
Published: (2025)
LibContinual: A Comprehensive Library towards Realistic Continual Learning
by: Li, Wenbin, et al.
Published: (2025)
by: Li, Wenbin, et al.
Published: (2025)
Adaptive Budget Allocation for Orthogonal-Subspace Adapter Tuning in LLMs Continual Learning
by: Wan, Zhiyi, et al.
Published: (2025)
by: Wan, Zhiyi, et al.
Published: (2025)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
by: Wang, Zi, et al.
Published: (2024)
by: Wang, Zi, et al.
Published: (2024)
Can LLMs Alleviate Catastrophic Forgetting in Graph Continual Learning? A Systematic Study
by: Cheng, Ziyang, et al.
Published: (2025)
by: Cheng, Ziyang, et al.
Published: (2025)
Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
by: Zhen, Shuai, et al.
Published: (2026)
by: Zhen, Shuai, et al.
Published: (2026)
Continuous Approximations for Improving Quantization Aware Training of LLMs
by: Li, He, et al.
Published: (2024)
by: Li, He, et al.
Published: (2024)
Similar Items
-
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
by: Yang, Zhen, et al.
Published: (2026) -
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024) -
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025) -
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026) -
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
by: Zhou, Huilin, et al.
Published: (2026)