WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Tianrong, Cao, Bochuan, Cao, Yuanpu, Lin, Lu, Mitra, Prasenjit, Chen, Jinghui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
by: Cao, Yuanpu, et al.
Published: (2024)
by: Cao, Yuanpu, et al.
Published: (2024)
PromptFix: Few-shot Backdoor Removal via Adversarial Prompt Tuning
by: Zhang, Tianrong, et al.
Published: (2024)
by: Zhang, Tianrong, et al.
Published: (2024)
Adversarially Robust Industrial Anomaly Detection Through Diffusion Model
by: Cao, Yuanpu, et al.
Published: (2024)
by: Cao, Yuanpu, et al.
Published: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
by: Cao, Yuanpu, et al.
Published: (2023)
by: Cao, Yuanpu, et al.
Published: (2023)
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
by: Lan, Yifan, et al.
Published: (2025)
by: Lan, Yifan, et al.
Published: (2025)
JoPA:Explaining Large Language Model's Generation via Joint Prompt Attribution
by: Chang, Yurui, et al.
Published: (2024)
by: Chang, Yurui, et al.
Published: (2024)
Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation
by: Chang, Yurui, et al.
Published: (2025)
by: Chang, Yurui, et al.
Published: (2025)
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
by: Lan, Yifan, et al.
Published: (2026)
by: Lan, Yifan, et al.
Published: (2026)
ForecastCompass: Guiding Agentic Forecasting with Adaptive Factor Memory
by: Chang, Yurui, et al.
Published: (2026)
by: Chang, Yurui, et al.
Published: (2026)
ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
by: Wang, Yujia, et al.
Published: (2025)
by: Wang, Yujia, et al.
Published: (2025)
Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
by: Liu, Zehao, et al.
Published: (2026)
by: Liu, Zehao, et al.
Published: (2026)
AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion models
by: Zeng, Yaopei, et al.
Published: (2024)
by: Zeng, Yaopei, et al.
Published: (2024)
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025)
by: Changjiang, Li, et al.
Published: (2025)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
by: Cao, Bochuan, et al.
Published: (2025)
by: Cao, Bochuan, et al.
Published: (2025)
OSNIP: Breaking the Privacy-Utility-Efficiency Trilemma in LLM Inference via Obfuscated Semantic Null Space
by: Cao, Zhiyuan, et al.
Published: (2026)
by: Cao, Zhiyuan, et al.
Published: (2026)
Transformer-Based Wildlife Species Classification from Daily Movement Trajectories
by: Irakoze, Obed, et al.
Published: (2026)
by: Irakoze, Obed, et al.
Published: (2026)
Automated Multi-Task Learning for Joint Disease Prediction on Electronic Health Records
by: Cui, Suhan, et al.
Published: (2024)
by: Cui, Suhan, et al.
Published: (2024)
Performance Anomaly Detection in Athletics: A Benchmarking System with Visual Analytics
by: Madukoma, Blessed, et al.
Published: (2026)
by: Madukoma, Blessed, et al.
Published: (2026)
Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization
by: Hu, Kai, et al.
Published: (2024)
by: Hu, Kai, et al.
Published: (2024)
PreFlect: From Retrospective to Prospective Reflection in Large Language Model Agents
by: Wang, Hanyu, et al.
Published: (2026)
by: Wang, Hanyu, et al.
Published: (2026)
Deep Generalized Schrödinger Bridges: From Image Generation to Solving Mean-Field Games
by: Liu, Guan-Horng, et al.
Published: (2024)
by: Liu, Guan-Horng, et al.
Published: (2024)
Defending Jailbreak Prompts via In-Context Adversarial Game
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game
by: Li, Lixing
Published: (2026)
by: Li, Lixing
Published: (2026)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
by: Chen, Wenyu, et al.
Published: (2026)
by: Chen, Wenyu, et al.
Published: (2026)
Good-Enough LLM Obfuscation (GELO)
by: Belikov, Anatoly, et al.
Published: (2026)
by: Belikov, Anatoly, et al.
Published: (2026)
Obfuscated Activations Bypass LLM Latent-Space Defenses
by: Bailey, Luke, et al.
Published: (2024)
by: Bailey, Luke, et al.
Published: (2024)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
by: Yin, Ziyi, et al.
Published: (2025)
by: Yin, Ziyi, et al.
Published: (2025)
AlignTree: Efficient Defense Against LLM Jailbreak Attacks
by: Goren, Gil, et al.
Published: (2025)
by: Goren, Gil, et al.
Published: (2025)
WildGraph: Realistic Graph-based Trajectory Generation for Wildlife
by: Al-Lawati, Ali, et al.
Published: (2024)
by: Al-Lawati, Ali, et al.
Published: (2024)
Semantic Captioning: Benchmark Dataset and Graph-Aware Few-Shot In-Context Learning for SQL2Text
by: Al-Lawati, Ali, et al.
Published: (2025)
by: Al-Lawati, Ali, et al.
Published: (2025)
Can Embedding Similarity Predict Cross-Lingual Transfer? A Systematic Study on African Languages
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
WildGEN: Long-horizon Trajectory Generation for Wildlife
by: Al-Lawati, Ali, et al.
Published: (2023)
by: Al-Lawati, Ali, et al.
Published: (2023)
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
by: Cao, Chentao, et al.
Published: (2025)
by: Cao, Chentao, et al.
Published: (2025)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
by: Zhang, Haonan, et al.
Published: (2026)
by: Zhang, Haonan, et al.
Published: (2026)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
by: Zhang, Chenyang, et al.
Published: (2026)
by: Zhang, Chenyang, et al.
Published: (2026)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
by: Zhang, Nan, et al.
Published: (2026)
by: Zhang, Nan, et al.
Published: (2026)
Similar Items
-
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023) -
TruthFlow: Truthful LLM Generation via Representation Flow Correction
by: Wang, Hanyu, et al.
Published: (2025) -
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
by: Cao, Yuanpu, et al.
Published: (2024) -
PromptFix: Few-shot Backdoor Removal via Adversarial Prompt Tuning
by: Zhang, Tianrong, et al.
Published: (2024) -
Adversarially Robust Industrial Anomaly Detection Through Diffusion Model
by: Cao, Yuanpu, et al.
Published: (2024)