Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Max, Liu, Derek, Zhang, Kai, Franco, Joshua, Liu, Haihao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreak Distillation: Renewable Safety Benchmarking
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Multilingual Jailbreak Challenges in Large Language Models
by: Deng, Yue, et al.
Published: (2023)
by: Deng, Yue, et al.
Published: (2023)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation
by: Huang, Chenyang, et al.
Published: (2025)
by: Huang, Chenyang, et al.
Published: (2025)
Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
Multilingual Knowledge Graph Completion via Efficient Multilingual Knowledge Sharing
by: Mao, Cunli, et al.
Published: (2025)
by: Mao, Cunli, et al.
Published: (2025)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
Why Agents Compromise Safety Under Pressure
by: Jiang, Hengle, et al.
Published: (2026)
by: Jiang, Hengle, et al.
Published: (2026)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
by: Huang, Caishuang, et al.
Published: (2024)
by: Huang, Caishuang, et al.
Published: (2024)
MPO: Multilingual Safety Alignment via Reward Gap Optimization
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Arctic-Embed 2.0: Multilingual Retrieval Without Compromise
by: Yu, Puxuan, et al.
Published: (2024)
by: Yu, Puxuan, et al.
Published: (2024)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
AfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languages
by: Raju, Joshua Sakthivel, et al.
Published: (2025)
by: Raju, Joshua Sakthivel, et al.
Published: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
by: Zhang, Yuanchi, et al.
Published: (2024)
by: Zhang, Yuanchi, et al.
Published: (2024)
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
by: Li, Xiaoxia, et al.
Published: (2024)
by: Li, Xiaoxia, et al.
Published: (2024)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
by: Gao, Lang, et al.
Published: (2024)
by: Gao, Lang, et al.
Published: (2024)
Enhancing Knowledge Distillation for LLMs with Response-Priming Prompting
by: Goyal, Vijay, et al.
Published: (2024)
by: Goyal, Vijay, et al.
Published: (2024)
TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language Models
by: Xu, Zhi, et al.
Published: (2026)
by: Xu, Zhi, et al.
Published: (2026)
Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
by: Liu, Lingyuan, et al.
Published: (2025)
by: Liu, Lingyuan, et al.
Published: (2025)
Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
by: Peng, Jingyu, et al.
Published: (2025)
by: Peng, Jingyu, et al.
Published: (2025)
Enhancing Low-Resource NMT with a Multilingual Encoder and Knowledge Distillation: A Case Study
by: Roy, Aniruddha, et al.
Published: (2024)
by: Roy, Aniruddha, et al.
Published: (2024)
Multilingual Knowledge Editing with Language-Agnostic Factual Neurons
by: Zhang, Xue, et al.
Published: (2024)
by: Zhang, Xue, et al.
Published: (2024)
Evolving Knowledge Distillation for Lightweight Neural Machine Translation
by: Zhang, Xuewen, et al.
Published: (2026)
by: Zhang, Xuewen, et al.
Published: (2026)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
Personalized LLM Response Generation with Parameterized Memory Injection
by: Zhang, Kai, et al.
Published: (2024)
by: Zhang, Kai, et al.
Published: (2024)
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
by: Song, Jialin, et al.
Published: (2026)
by: Song, Jialin, et al.
Published: (2026)
Tracing Multilingual Factual Knowledge Acquisition in Pretraining
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
by: Liu, Jiarui, et al.
Published: (2026)
by: Liu, Jiarui, et al.
Published: (2026)
Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
by: Sriratanawilai, Sukrit, et al.
Published: (2025)
by: Sriratanawilai, Sukrit, et al.
Published: (2025)
Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models
by: Rao, Jun, et al.
Published: (2024)
by: Rao, Jun, et al.
Published: (2024)
Generating Place-Based Compromises Between Two Points of View
by: Bhattacharyya, Sumanta, et al.
Published: (2026)
by: Bhattacharyya, Sumanta, et al.
Published: (2026)
Distillation for Multilingual Information Retrieval
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
Similar Items
-
Jailbreak Distillation: Renewable Safety Benchmarking
by: Zhang, Jingyu, et al.
Published: (2025) -
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026) -
Multilingual Jailbreak Challenges in Large Language Models
by: Deng, Yue, et al.
Published: (2023) -
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
by: Lu, Xiaoya, et al.
Published: (2025) -
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation
by: Huang, Chenyang, et al.
Published: (2025)