Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Yixin, Georgopoulos, Markos, Cevher, Volkan, Chrysos, Grigorios G. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Certified Robustness Under Bounded Levenshtein Distance
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
Single-pass Detection of Jailbreaking Input in Large Language Models
by: Candogan, Leyla Naz, et al.
Published: (2025)
by: Candogan, Leyla Naz, et al.
Published: (2025)
Multilinear Operator Networks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
Revisiting Character-level Adversarial Attacks for Language Models
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
by: Afzal, Arshia, et al.
Published: (2024)
by: Afzal, Arshia, et al.
Published: (2024)
Robust NAS under adversarial training: benchmark, theory, and beyond
by: Wu, Yongtao, et al.
Published: (2024)
by: Wu, Yongtao, et al.
Published: (2024)
Efficient local linearity regularization to overcome catastrophic overfitting
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
Going beyond Compositions, DDPMs Can Produce Zero-Shot Interpolations
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
by: Wu, Yongtao, et al.
Published: (2025)
by: Wu, Yongtao, et al.
Published: (2025)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
by: Xiong, Zheyang, et al.
Published: (2024)
by: Xiong, Zheyang, et al.
Published: (2024)
The Last Mile to Supervised Performance: Semi-Supervised Domain Adaptation for Semantic Segmentation
by: Morales-Brotons, Daniel, et al.
Published: (2024)
by: Morales-Brotons, Daniel, et al.
Published: (2024)
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026)
by: Misra, Diganta, et al.
Published: (2026)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
Membership Inference Attacks against Large Vision-Language Models
by: Li, Zhan, et al.
Published: (2024)
by: Li, Zhan, et al.
Published: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions
by: Tan, Cheng, et al.
Published: (2024)
by: Tan, Cheng, et al.
Published: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Learning to Remove Cuts in Integer Linear Programming
by: Puigdemont, Pol, et al.
Published: (2024)
by: Puigdemont, Pol, et al.
Published: (2024)
Boosting Jailbreak Attack with Momentum
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
by: Cheng, Yixin, et al.
Published: (2025)
by: Cheng, Yixin, et al.
Published: (2025)
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
by: Mura, Raffaele, et al.
Published: (2025)
by: Mura, Raffaele, et al.
Published: (2025)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
by: Fang, Zhicheng, et al.
Published: (2026)
by: Fang, Zhicheng, et al.
Published: (2026)
Jailbreaking with Universal Multi-Prompts
by: Hsu, Yu-Ling, et al.
Published: (2025)
by: Hsu, Yu-Ling, et al.
Published: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
by: Cheng, Wenhua, et al.
Published: (2023)
by: Cheng, Wenhua, et al.
Published: (2023)
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024)
by: Lee, Isack, et al.
Published: (2024)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
Generalization of Scaled Deep ResNets in the Mean-Field Regime
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Hadamard product in deep learning: Introduction, Advances and Challenges
by: Chrysos, Grigorios G, et al.
Published: (2025)
by: Chrysos, Grigorios G, et al.
Published: (2025)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
by: Wei, Zeming, et al.
Published: (2023)
by: Wei, Zeming, et al.
Published: (2023)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
by: Jiang, Weisen, et al.
Published: (2025)
by: Jiang, Weisen, et al.
Published: (2025)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
by: Lin, Shi, et al.
Published: (2024)
by: Lin, Shi, et al.
Published: (2024)
Similar Items
-
Certified Robustness Under Bounded Levenshtein Distance
by: Rocamora, Elias Abad, et al.
Published: (2025) -
Single-pass Detection of Jailbreaking Input in Large Language Models
by: Candogan, Leyla Naz, et al.
Published: (2025) -
Multilinear Operator Networks
by: Cheng, Yixin, et al.
Published: (2024) -
Revisiting Character-level Adversarial Attacks for Language Models
by: Rocamora, Elias Abad, et al.
Published: (2024) -
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
by: Afzal, Arshia, et al.
Published: (2024)