Automating Deception: Scalable Multi-Turn LLM Jailbreaks
Fuente:
arXiv
Saved in:
| Main Authors: | Kumarappan, Adarsh, Mujoo, Ananya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
by: He, Zhida, et al.
Published: (2026)
by: He, Zhida, et al.
Published: (2026)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
by: Karnam, Meghana, et al.
Published: (2026)
by: Karnam, Meghana, et al.
Published: (2026)
TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
by: Xiong, Xiqiao, et al.
Published: (2025)
by: Xiong, Xiqiao, et al.
Published: (2025)
LeanAgent: Lifelong Learning for Formal Theorem Proving
by: Kumarappan, Adarsh, et al.
Published: (2024)
by: Kumarappan, Adarsh, et al.
Published: (2024)
Deceptive Exploration in Multi-armed Bandits
by: Vurankaya, I. Arda, et al.
Published: (2025)
by: Vurankaya, I. Arda, et al.
Published: (2025)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
by: Agnihotri, Rudransh, et al.
Published: (2025)
by: Agnihotri, Rudransh, et al.
Published: (2025)
Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores
by: Panda, Shevya, et al.
Published: (2026)
by: Panda, Shevya, et al.
Published: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
by: Liu, Licheng, et al.
Published: (2025)
by: Liu, Licheng, et al.
Published: (2025)
Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
Gradient Multi-Normalization for Stateless and Scalable LLM Training
by: Scetbon, Meyer, et al.
Published: (2025)
by: Scetbon, Meyer, et al.
Published: (2025)
Deception at Scale: Deceptive Designs in 1K LLM-Generated Ecommerce Components
by: Chen, Ziwei, et al.
Published: (2025)
by: Chen, Ziwei, et al.
Published: (2025)
Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry
by: Zhang, Guoxi, et al.
Published: (2026)
by: Zhang, Guoxi, et al.
Published: (2026)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Detecting Jailbreak Attempts in Clinical Training LLMs Through Automated Linguistic Feature Extraction
by: Nguyen, Tri, et al.
Published: (2026)
by: Nguyen, Tri, et al.
Published: (2026)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning
by: Jali, Neharika, et al.
Published: (2026)
by: Jali, Neharika, et al.
Published: (2026)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
by: Dong, Yingkai, et al.
Published: (2024)
by: Dong, Yingkai, et al.
Published: (2024)
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
by: Djuhera, Aladin, et al.
Published: (2026)
by: Djuhera, Aladin, et al.
Published: (2026)
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
by: Chai, Jiajun, et al.
Published: (2025)
by: Chai, Jiajun, et al.
Published: (2025)
LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
by: Zhang, Haonan, et al.
Published: (2026)
by: Zhang, Haonan, et al.
Published: (2026)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Mitigating Conversational Inertia in Multi-Turn Agents
by: Wan, Yang, et al.
Published: (2026)
by: Wan, Yang, et al.
Published: (2026)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
Turning LLM Activations Quantization-Friendly
by: Czakó, Patrik, et al.
Published: (2025)
by: Czakó, Patrik, et al.
Published: (2025)
Information-theoretic Distinctions Between Deception and Confusion
by: Young, Robin
Published: (2025)
by: Young, Robin
Published: (2025)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
by: Zolfaghari, Vahideh
Published: (2026)
by: Zolfaghari, Vahideh
Published: (2026)
LLM-AR: LLM-powered Automated Reasoning Framework
by: Chen, Rick, et al.
Published: (2025)
by: Chen, Rick, et al.
Published: (2025)
Adversarial Reasoning at Jailbreaking Time
by: Sabbaghi, Mahdi, et al.
Published: (2025)
by: Sabbaghi, Mahdi, et al.
Published: (2025)
When Truthful Representations Flip Under Deceptive Instructions?
by: Long, Xianxuan, et al.
Published: (2025)
by: Long, Xianxuan, et al.
Published: (2025)
POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue
by: Cao, Ruike, et al.
Published: (2026)
by: Cao, Ruike, et al.
Published: (2026)
Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
by: Wei, Chenxing, et al.
Published: (2026)
by: Wei, Chenxing, et al.
Published: (2026)
Ethical and Scalable Automation: A Governance and Compliance Framework for Business Applications
by: Lin, Haocheng
Published: (2024)
by: Lin, Haocheng
Published: (2024)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Similar Items
-
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
by: Kumarappan, Adarsh, et al.
Published: (2026) -
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
by: Kumarappan, Adarsh, et al.
Published: (2025) -
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
by: He, Zhida, et al.
Published: (2026) -
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025) -
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
by: Karnam, Meghana, et al.
Published: (2026)