Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Ruohao, Oroojlooy, Afshin, Sridhar, Roshan, Ballesteros, Miguel, Ritter, Alan, Roth, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
by: Guo, Ruohao, et al.
Published: (2025)
by: Guo, Ruohao, et al.
Published: (2025)
Evaluating NL2SQL via SQL2NL
by: Safarzadeh, Mohammadtaher, et al.
Published: (2025)
by: Safarzadeh, Mohammadtaher, et al.
Published: (2025)
Investigating and Alleviating Harm Amplification in LLM Interactions
by: Guo, Ruohao, et al.
Published: (2026)
by: Guo, Ruohao, et al.
Published: (2026)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
by: Guo, Ruohao, et al.
Published: (2023)
by: Guo, Ruohao, et al.
Published: (2023)
Red Teaming Language Models for Processing Contradictory Dialogues
by: Wen, Xiaofei, et al.
Published: (2024)
by: Wen, Xiaofei, et al.
Published: (2024)
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
by: Beutel, Alex, et al.
Published: (2024)
by: Beutel, Alex, et al.
Published: (2024)
Language Models can Self-Improve at State-Value Estimation for Better Search
by: Mendes, Ethan, et al.
Published: (2025)
by: Mendes, Ethan, et al.
Published: (2025)
Agentic Reinforced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)
by: Yu, Xiaodong, et al.
Published: (2023)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
by: Khadangi, Afshin, et al.
Published: (2025)
by: Khadangi, Afshin, et al.
Published: (2025)
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
by: Guo, Yiran, et al.
Published: (2025)
by: Guo, Yiran, et al.
Published: (2025)
MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
by: Yan, Siyu, et al.
Published: (2025)
by: Yan, Siyu, et al.
Published: (2025)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
by: Park, Jungsoo, et al.
Published: (2026)
by: Park, Jungsoo, et al.
Published: (2026)
Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
by: Ding, Jiale, et al.
Published: (2025)
by: Ding, Jiale, et al.
Published: (2025)
Anticipatory Evaluation of Language Models
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
by: Cheng, Gang, et al.
Published: (2025)
by: Cheng, Gang, et al.
Published: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
by: Walder, Christian, et al.
Published: (2025)
by: Walder, Christian, et al.
Published: (2025)
MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning
by: Wang, Hongjun, et al.
Published: (2026)
by: Wang, Hongjun, et al.
Published: (2026)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
When Vision-Language Models Judge Without Seeing: Exposing Informativeness Bias
by: Zou, Xiaohan, et al.
Published: (2026)
by: Zou, Xiaohan, et al.
Published: (2026)
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
by: Naous, Tarek, et al.
Published: (2023)
by: Naous, Tarek, et al.
Published: (2023)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
DCPO: Dynamic Clipping Policy Optimization
by: Yang, Shihui, et al.
Published: (2025)
by: Yang, Shihui, et al.
Published: (2025)
SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks
by: Safarzadeh, Mohammadtaher, et al.
Published: (2026)
by: Safarzadeh, Mohammadtaher, et al.
Published: (2026)
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025)
by: Nakash, Itay, et al.
Published: (2025)
An LLM Feature-based Framework for Dialogue Constructiveness Assessment
by: Zhou, Lexin, et al.
Published: (2024)
by: Zhou, Lexin, et al.
Published: (2024)
Reinforcement Learning for Personalized Dialogue Management
by: Hengst, Floris den, et al.
Published: (2019)
by: Hengst, Floris den, et al.
Published: (2019)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
Fibration Policy Optimization
by: Li, Chang, et al.
Published: (2026)
by: Li, Chang, et al.
Published: (2026)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
by: Sharma, Mrinank, et al.
Published: (2025)
by: Sharma, Mrinank, et al.
Published: (2025)
Anecdoctoring: Automated Red-Teaming Across Language and Place
by: Cuevas, Alejandro, et al.
Published: (2025)
by: Cuevas, Alejandro, et al.
Published: (2025)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
by: Yao, Jiashu, et al.
Published: (2026)
by: Yao, Jiashu, et al.
Published: (2026)
Similar Items
-
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
by: Guo, Ruohao, et al.
Published: (2025) -
Evaluating NL2SQL via SQL2NL
by: Safarzadeh, Mohammadtaher, et al.
Published: (2025) -
Investigating and Alleviating Harm Amplification in LLM Interactions
by: Guo, Ruohao, et al.
Published: (2026) -
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
by: Guo, Ruohao, et al.
Published: (2023) -
Red Teaming Language Models for Processing Contradictory Dialogues
by: Wen, Xiaofei, et al.
Published: (2024)