Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yao, Song, Zeen, Qiang, Wenwen, Wu, Fengge, Zhou, Shuyi, Zheng, Changwen, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024)
by: Zhang, Congzhi, et al.
Published: (2024)
Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024)
by: Zhang, Congzhi, et al.
Published: (2024)
Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective
by: Song, Zeen, et al.
Published: (2024)
by: Song, Zeen, et al.
Published: (2024)
On the Discriminability of Self-Supervised Representation Learning
by: Song, Zeen, et al.
Published: (2024)
by: Song, Zeen, et al.
Published: (2024)
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
On the Generalization and Causal Explanation in Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2024)
by: Qiang, Wenwen, et al.
Published: (2024)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
by: Guo, Peizheng, et al.
Published: (2025)
by: Guo, Peizheng, et al.
Published: (2025)
Learning Invariant Causal Mechanism from Vision-Language Models
by: Song, Zeen, et al.
Published: (2024)
by: Song, Zeen, et al.
Published: (2024)
BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain Abstraction
by: Li, Jiangmeng, et al.
Published: (2024)
by: Li, Jiangmeng, et al.
Published: (2024)
Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
by: Weng, Zixuan, et al.
Published: (2025)
by: Weng, Zixuan, et al.
Published: (2025)
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
by: Song, Ruike, et al.
Published: (2025)
by: Song, Ruike, et al.
Published: (2025)
Adaptive Uncertainty-Aware Tree Search for Robust Reasoning
by: Song, Zeen, et al.
Published: (2026)
by: Song, Zeen, et al.
Published: (2026)
Manifold Constraint Regularization for Remote Sensing Image Generation
by: Su, Xingzhe, et al.
Published: (2023)
by: Su, Xingzhe, et al.
Published: (2023)
Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
From Shallow to Deep: Pinning Semantic Intent via Causal GRPO
by: Zhou, Shuyi, et al.
Published: (2026)
by: Zhou, Shuyi, et al.
Published: (2026)
Unbiased Image Synthesis via Manifold Guidance in Diffusion Models
by: Su, Xingzhe, et al.
Published: (2023)
by: Su, Xingzhe, et al.
Published: (2023)
Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized Approach
by: Gao, Hang, et al.
Published: (2024)
by: Gao, Hang, et al.
Published: (2024)
Intriguing Property and Counterfactual Explanation of GAN for Remote Sensing Image Generation
by: Su, Xingzhe, et al.
Published: (2023)
by: Su, Xingzhe, et al.
Published: (2023)
Dialogue Injection Attack: Jailbreaking LLMs through Context Manipulation
by: Meng, Wenlong, et al.
Published: (2025)
by: Meng, Wenlong, et al.
Published: (2025)
Beyond All-to-All: Causal-Aligned Transformer with Dynamic Structure Learning for Multivariate Time Series Forecasting
by: Zhang, Xingyu, et al.
Published: (2025)
by: Zhang, Xingyu, et al.
Published: (2025)
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
by: Chang, Wenhan, et al.
Published: (2025)
by: Chang, Wenhan, et al.
Published: (2025)
Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
by: Lin, Yuping, et al.
Published: (2024)
by: Lin, Yuping, et al.
Published: (2024)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
Closing the Loop: A Control-Theoretic Framework for Provably Stable Time Series Forecasting with LLMs
by: Zhang, Xingyu, et al.
Published: (2026)
by: Zhang, Xingyu, et al.
Published: (2026)
On the Out-of-Distribution Generalization of Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2025)
by: Qiang, Wenwen, et al.
Published: (2025)
Reward Model Generalization for Compute-Aware Test-Time Reasoning
by: Song, Zeen, et al.
Published: (2025)
by: Song, Zeen, et al.
Published: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
by: Zhou, Andy, et al.
Published: (2024)
by: Zhou, Andy, et al.
Published: (2024)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
Group Causal Policy Optimization for Post-Training Large Language Models
by: Gu, Ziyin, et al.
Published: (2025)
by: Gu, Ziyin, et al.
Published: (2025)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
by: Wang, Zijun, et al.
Published: (2024)
by: Wang, Zijun, et al.
Published: (2024)
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
Interventional Imbalanced Multi-Modal Representation Learning via $β$-Generalization Front-Door Criterion
by: Li, Yi, et al.
Published: (2024)
by: Li, Yi, et al.
Published: (2024)
Defending LLMs against Jailbreaking Attacks via Backtranslation
by: Wang, Yihan, et al.
Published: (2024)
by: Wang, Yihan, et al.
Published: (2024)
Causal Prompt Calibration Guided Segment Anything Model for Open-Vocabulary Multi-Entity Segmentation
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
by: Pu, Rui, et al.
Published: (2024)
by: Pu, Rui, et al.
Published: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Jailbreaking Large Language Models with Morality Attacks
by: Su, Ying, et al.
Published: (2026)
by: Su, Ying, et al.
Published: (2026)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Similar Items
-
Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024) -
Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024) -
Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective
by: Song, Zeen, et al.
Published: (2024) -
On the Discriminability of Self-Supervised Representation Learning
by: Song, Zeen, et al.
Published: (2024) -
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
by: Wang, Jingyao, et al.
Published: (2025)