Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ming, Chen, Pei, Zhang, Zhenhao, Yang, Tao, Zhang, Xinyang, Li, Han, Cao, Tianyu, Zeng, Ming, Wu, Zhuofeng, Jiang, Meng, Li, Huasheng, Li, Lihong, Yin, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Mitigating Conversational Inertia in Multi-Turn Agents
by: Wan, Yang, et al.
Published: (2026)
by: Wan, Yang, et al.
Published: (2026)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
by: Tan, Chuyi, et al.
Published: (2025)
by: Tan, Chuyi, et al.
Published: (2025)
UniConv: Unifying Retrieval and Response Generation for Large Language Models in Conversations
by: Mo, Fengran, et al.
Published: (2025)
by: Mo, Fengran, et al.
Published: (2025)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
by: Zhai, Skylar, et al.
Published: (2026)
by: Zhai, Skylar, et al.
Published: (2026)
Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates
by: Dang, Hy, et al.
Published: (2025)
by: Dang, Hy, et al.
Published: (2025)
Rapid Computation of the Plasma Dispersion Function: Rational and Multi-pole Approximation, and Improved Accuracy
by: Xie, Huasheng
Published: (2024)
by: Xie, Huasheng
Published: (2024)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
by: Zhang, Jiazheng, et al.
Published: (2026)
by: Zhang, Jiazheng, et al.
Published: (2026)
Deeper with Riemannian Geometry: Overcoming Oversmoothing and Oversquashing for Graph Foundation Models
by: Sun, Li, et al.
Published: (2025)
by: Sun, Li, et al.
Published: (2025)
END: Early Noise Dropping for Efficient and Effective Context Denoising
by: Jin, Hongye, et al.
Published: (2025)
by: Jin, Hongye, et al.
Published: (2025)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
by: Bai, Yuyang, et al.
Published: (2026)
by: Bai, Yuyang, et al.
Published: (2026)
From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
by: Gao, Jiaxuan, et al.
Published: (2026)
by: Gao, Jiaxuan, et al.
Published: (2026)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
by: Li, Gaotang, et al.
Published: (2026)
by: Li, Gaotang, et al.
Published: (2026)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
by: Lei, Xuanyu, et al.
Published: (2025)
by: Lei, Xuanyu, et al.
Published: (2025)
A Verifiable Dynamic (t,n) Threshold Quantum Secret Sharing Protocol with Authentication
by: Jian‐Qiu Li, et al.
Published: (2025)
by: Jian‐Qiu Li, et al.
Published: (2025)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Mitigating Distribution Shift in Model-based Offline RL via Shifts-aware Reward Learning
by: Luo, Wang, et al.
Published: (2024)
by: Luo, Wang, et al.
Published: (2024)
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation
by: Zong, Haotian, et al.
Published: (2026)
by: Zong, Haotian, et al.
Published: (2026)
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
by: Li, Junbo, et al.
Published: (2025)
by: Li, Junbo, et al.
Published: (2025)
XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
by: Chen, Dian, et al.
Published: (2025)
by: Chen, Dian, et al.
Published: (2025)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
What Is the Minimum Number of Parameters Required to Represent Solutions of the Grad-Shafranov Equation?
by: Xie, Huasheng, et al.
Published: (2026)
by: Xie, Huasheng, et al.
Published: (2026)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
by: Zhang, Qingru, et al.
Published: (2025)
by: Zhang, Qingru, et al.
Published: (2025)
The correlation between nativelike selection and prototypicality: a multilingual onomasiological case study using semantic embedding
by: Zhang, Huasheng
Published: (2024)
by: Zhang, Huasheng
Published: (2024)
What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis
by: Li, Xirui, et al.
Published: (2026)
by: Li, Xirui, et al.
Published: (2026)
Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents
by: Chuang, Yun-Shiuan, et al.
Published: (2026)
by: Chuang, Yun-Shiuan, et al.
Published: (2026)
MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution Reasoning
by: Tang, Lingxiao, et al.
Published: (2026)
by: Tang, Lingxiao, et al.
Published: (2026)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
by: Zhang, Zijing, et al.
Published: (2025)
by: Zhang, Zijing, et al.
Published: (2025)
IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination
by: Li, Yuanshuai, et al.
Published: (2026)
by: Li, Yuanshuai, et al.
Published: (2026)
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
by: Pavlenko, Kirill, et al.
Published: (2026)
by: Pavlenko, Kirill, et al.
Published: (2026)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting
by: Li, Renda, et al.
Published: (2025)
by: Li, Renda, et al.
Published: (2025)
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
by: Fan, Mingyuan, et al.
Published: (2026)
by: Fan, Mingyuan, et al.
Published: (2026)
Similar Items
-
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026) -
Mitigating Conversational Inertia in Multi-Turn Agents
by: Wan, Yang, et al.
Published: (2026) -
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
by: Zhang, Xin, et al.
Published: (2026) -
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
by: Su, Yi, et al.
Published: (2025) -
Diagnosing and Mitigating System Bias in Self-Rewarding RL
by: Tan, Chuyi, et al.
Published: (2025)