The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Mingrui, Zhang, Sixiao, Long, Cheng, Lam, Kwok Yan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
by: Liu, Mingrui, et al.
Published: (2026)
by: Liu, Mingrui, et al.
Published: (2026)
Mask-based Membership Inference Attacks for Retrieval-Augmented Generation
by: Liu, Mingrui, et al.
Published: (2024)
by: Liu, Mingrui, et al.
Published: (2024)
Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
by: Liu, Mingrui, et al.
Published: (2025)
by: Liu, Mingrui, et al.
Published: (2025)
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
by: Wang, Yanting, et al.
Published: (2025)
by: Wang, Yanting, et al.
Published: (2025)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
by: Xie, Zhixin, et al.
Published: (2026)
by: Xie, Zhixin, et al.
Published: (2026)
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
A Comparative Study of Fuzzers and Static Analysis Tools for Finding Memory Unsafety in C and C++
by: Hassler, Keno, et al.
Published: (2025)
by: Hassler, Keno, et al.
Published: (2025)
Neural Trojans
by: Liu, Yuntao, et al.
Published: (2017)
by: Liu, Yuntao, et al.
Published: (2017)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
by: Lv, Lijia, et al.
Published: (2024)
by: Lv, Lijia, et al.
Published: (2024)
TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models
by: Zhang, Chaoshuo, et al.
Published: (2026)
by: Zhang, Chaoshuo, et al.
Published: (2026)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
by: Wang, Xinkai, et al.
Published: (2025)
by: Wang, Xinkai, et al.
Published: (2025)
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
by: Faruque, Md Omar, et al.
Published: (2024)
by: Faruque, Md Omar, et al.
Published: (2024)
Beyond Fixed and Dynamic Prompts: Embedded Jailbreak Templates for Advancing LLM Security
by: Kim, Hajun, et al.
Published: (2025)
by: Kim, Hajun, et al.
Published: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
by: Cui, Tiehan, et al.
Published: (2025)
by: Cui, Tiehan, et al.
Published: (2025)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
by: Shen, Qingchao, et al.
Published: (2026)
by: Shen, Qingchao, et al.
Published: (2026)
TrojanLoC: LLM-based Framework for RTL Trojan Localization
by: Xiao, Weihua, et al.
Published: (2025)
by: Xiao, Weihua, et al.
Published: (2025)
The Philosopher's Stone: Trojaning Plugins of Large Language Models
by: Dong, Tian, et al.
Published: (2023)
by: Dong, Tian, et al.
Published: (2023)
CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
by: Nahian, Mohaiminul Al, et al.
Published: (2025)
by: Nahian, Mohaiminul Al, et al.
Published: (2025)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
by: Ahn, Yelim, et al.
Published: (2025)
by: Ahn, Yelim, et al.
Published: (2025)
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
by: Jeong, Joonhyun, et al.
Published: (2025)
by: Jeong, Joonhyun, et al.
Published: (2025)
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
by: Yan, Yu, et al.
Published: (2026)
by: Yan, Yu, et al.
Published: (2026)
Alphabet Index Mapping: Jailbreaking LLMs through Semantic Dissimilarity
by: Husain, Bilal Saleh
Published: (2025)
by: Husain, Bilal Saleh
Published: (2025)
HeisenTrojans: They Are Not There Until They Are Triggered
by: Mavurapu, Akshita Reddy, et al.
Published: (2023)
by: Mavurapu, Akshita Reddy, et al.
Published: (2023)
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
by: Liu, Fazhong, et al.
Published: (2026)
by: Liu, Fazhong, et al.
Published: (2026)
Watermarking Recommender Systems
by: Zhang, Sixiao, et al.
Published: (2024)
by: Zhang, Sixiao, et al.
Published: (2024)
ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
TrojanForge: Generating Adversarial Hardware Trojan Examples Using Reinforcement Learning
by: Sarihi, Amin, et al.
Published: (2024)
by: Sarihi, Amin, et al.
Published: (2024)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
by: Sel, Bilgehan, et al.
Published: (2026)
by: Sel, Bilgehan, et al.
Published: (2026)
Unlearnable Examples Detection via Iterative Filtering
by: Yu, Yi, et al.
Published: (2024)
by: Yu, Yi, et al.
Published: (2024)
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
by: Zou, Quanchen, et al.
Published: (2026)
by: Zou, Quanchen, et al.
Published: (2026)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Obfuscating IoT Device Scanning Activity via Adversarial Example Generation
by: Li, Haocong, et al.
Published: (2024)
by: Li, Haocong, et al.
Published: (2024)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
by: Liu, Yupei, et al.
Published: (2025)
by: Liu, Yupei, et al.
Published: (2025)
Efficient Privacy-Preserving Retrieval Augmented Generation with Distance-Preserving Encryption
by: Ye, Huanyi, et al.
Published: (2026)
by: Ye, Huanyi, et al.
Published: (2026)
Guaranteeing Data Privacy in Federated Unlearning with Dynamic User Participation
by: Liu, Ziyao, et al.
Published: (2024)
by: Liu, Ziyao, et al.
Published: (2024)
Similar Items
-
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
by: Liu, Mingrui, et al.
Published: (2026) -
Mask-based Membership Inference Attacks for Retrieval-Augmented Generation
by: Liu, Mingrui, et al.
Published: (2024) -
Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
by: Liu, Mingrui, et al.
Published: (2025) -
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
by: Wang, Yanting, et al.
Published: (2025) -
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
by: Xie, Zhixin, et al.
Published: (2026)