RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Tuan T., Le, John, Vu, Thai T., Susilo, Willy, Cooper, Heath |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024)
by: Zhou, Yukai, et al.
Published: (2024)
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
by: Nguyen, Tuc, et al.
Published: (2026)
by: Nguyen, Tuc, et al.
Published: (2026)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
Generalizability of Mixture of Domain-Specific Adapters from the Lens of Signed Weight Directions and its Application to Effective Model Pruning
by: Nguyen, Tuc, et al.
Published: (2024)
by: Nguyen, Tuc, et al.
Published: (2024)
Adapters Mixup: Mixing Parameter-Efficient Adapters to Enhance the Adversarial Robustness of Fine-tuned Pre-trained Text Classifiers
by: Nguyen, Tuc, et al.
Published: (2024)
by: Nguyen, Tuc, et al.
Published: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Mitigating Jailbreaks with Intent-Aware LLMs
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Silenced Biases: The Dark Side LLMs Learned to Refuse
by: Himelstein, Rom, et al.
Published: (2025)
by: Himelstein, Rom, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
by: Maskey, Utsav, et al.
Published: (2026)
by: Maskey, Utsav, et al.
Published: (2026)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
by: Nguyen, Cong-Duy, et al.
Published: (2023)
by: Nguyen, Cong-Duy, et al.
Published: (2023)
Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verification
by: Nguyen, Tuc, et al.
Published: (2025)
by: Nguyen, Tuc, et al.
Published: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
by: Shrivastava, Aryan, et al.
Published: (2025)
by: Shrivastava, Aryan, et al.
Published: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
by: Pham, Viet, et al.
Published: (2025)
by: Pham, Viet, et al.
Published: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
TOPFORMER: Topology-Aware Authorship Attribution of Deepfake Texts with Diverse Writing Styles
by: Uchendu, Adaku, et al.
Published: (2023)
by: Uchendu, Adaku, et al.
Published: (2023)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers
by: Nguyen, Viet-Anh, et al.
Published: (2025)
by: Nguyen, Viet-Anh, et al.
Published: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
by: von Recum, Alexander, et al.
Published: (2024)
by: von Recum, Alexander, et al.
Published: (2024)
Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
by: Ho, Tuan Vu, et al.
Published: (2025)
by: Ho, Tuan Vu, et al.
Published: (2025)
Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
by: Zhao, Yinzhi, et al.
Published: (2026)
by: Zhao, Yinzhi, et al.
Published: (2026)
NoMatterXAI: Generating "No Matter What" Alterfactual Examples for Explaining Black-Box Text Classification Models
by: Nguyen, Tuc, et al.
Published: (2024)
by: Nguyen, Tuc, et al.
Published: (2024)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
by: Mavi, John, et al.
Published: (2025)
by: Mavi, John, et al.
Published: (2025)
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling
by: Yang, Jeff, et al.
Published: (2025)
by: Yang, Jeff, et al.
Published: (2025)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
by: Song, Maojia, et al.
Published: (2024)
by: Song, Maojia, et al.
Published: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
by: Hendria, Willy Fitra
Published: (2026)
by: Hendria, Willy Fitra
Published: (2026)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
ViMultiChoice: Toward a Method That Gives Explanation for Multiple-Choice Reading Comprehension in Vietnamese
by: Cao, Trung Tien, et al.
Published: (2026)
by: Cao, Trung Tien, et al.
Published: (2026)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
by: Dugan, Liam, et al.
Published: (2024)
by: Dugan, Liam, et al.
Published: (2024)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
Topological Data Analysis Applications in Natural Language Processing: A Survey
by: Uchendu, Adaku, et al.
Published: (2024)
by: Uchendu, Adaku, et al.
Published: (2024)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
by: Zhu, Runchuan, et al.
Published: (2025)
by: Zhu, Runchuan, et al.
Published: (2025)
Similar Items
-
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024) -
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
by: Nguyen, Tuc, et al.
Published: (2026) -
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024) -
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
by: Pan, Wenbo, et al.
Published: (2025) -
Generalizability of Mixture of Domain-Specific Adapters from the Lens of Signed Weight Directions and its Application to Effective Model Pruning
by: Nguyen, Tuc, et al.
Published: (2024)