Saved in:
| Main Authors: | Zhang, Xikai, Li, Yongzhi, Xiao, Likang, Zhang, Yingze, Cheng, Yanhua, Chen, Quan, Jiang, Peng, Wu, Wenjun, Liu, Liu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.20256 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
by: Zhang, Xikai, et al.
Published: (2025)
by: Zhang, Xikai, et al.
Published: (2025)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
by: Cui, Yitong, et al.
Published: (2025)
by: Cui, Yitong, et al.
Published: (2025)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
by: Ren, Yanwei, et al.
Published: (2026)
by: Ren, Yanwei, et al.
Published: (2026)
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
by: Zhang, Haotian, et al.
Published: (2025)
by: Zhang, Haotian, et al.
Published: (2025)
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
by: Sun, Haoran, et al.
Published: (2026)
by: Sun, Haoran, et al.
Published: (2026)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
by: Xu, Yifan, et al.
Published: (2025)
by: Xu, Yifan, et al.
Published: (2025)
MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control
by: Liu, Zifan, et al.
Published: (2024)
by: Liu, Zifan, et al.
Published: (2024)
Enhancing Collaborative Semantics of Language Model-Driven Recommendations via Graph-Aware Learning
by: Guan, Zhong, et al.
Published: (2024)
by: Guan, Zhong, et al.
Published: (2024)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
by: Liu, Bingshuai, et al.
Published: (2025)
by: Liu, Bingshuai, et al.
Published: (2025)
Large Language Models as Amortized Pareto-Front Generators for Constrained Bi-Objective Convex Optimization
by: Xu, Peipei, et al.
Published: (2026)
by: Xu, Peipei, et al.
Published: (2026)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems
by: Liu, Kun, et al.
Published: (2026)
by: Liu, Kun, et al.
Published: (2026)
Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
by: Yuan, Wenzhen, et al.
Published: (2025)
by: Yuan, Wenzhen, et al.
Published: (2025)
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
by: Yao, Zhengjian, et al.
Published: (2026)
by: Yao, Zhengjian, et al.
Published: (2026)
Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
by: You, Yuyang, et al.
Published: (2026)
by: You, Yuyang, et al.
Published: (2026)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
by: Liu, Shaoteng, et al.
Published: (2024)
by: Liu, Shaoteng, et al.
Published: (2024)
STEMO: Early Spatio-temporal Forecasting with Multi-Objective Reinforcement Learning
by: Shao, Wei, et al.
Published: (2024)
by: Shao, Wei, et al.
Published: (2024)
Reinforcement Learning from User Feedback
by: Han, Eric, et al.
Published: (2025)
by: Han, Eric, et al.
Published: (2025)
Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
by: Deng, Zehao, et al.
Published: (2025)
by: Deng, Zehao, et al.
Published: (2025)
Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
by: Zhen, Shuai, et al.
Published: (2026)
by: Zhen, Shuai, et al.
Published: (2026)
Robust Evolutionary Multi-Objective Network Architecture Search for Reinforcement Learning (EMNAS-RL)
by: Adde, Nihal Acharya, et al.
Published: (2025)
by: Adde, Nihal Acharya, et al.
Published: (2025)
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
by: Lai, Hanyu, et al.
Published: (2025)
by: Lai, Hanyu, et al.
Published: (2025)
Adventurer: Exploration with BiGAN for Deep Reinforcement Learning
by: Liu, Yongshuai, et al.
Published: (2025)
by: Liu, Yongshuai, et al.
Published: (2025)
Balancing Multiple Objectives in Urban Traffic Control with Reinforcement Learning from AI Feedback
by: Zhao, Chenyang, et al.
Published: (2026)
by: Zhao, Chenyang, et al.
Published: (2026)
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
by: Peng, Xiyue, et al.
Published: (2024)
by: Peng, Xiyue, et al.
Published: (2024)
Hierarchical Semantic RL: Tackling the Problem of Dynamic Action Space for RL-based Recommendations
by: Wang, Minmao, et al.
Published: (2025)
by: Wang, Minmao, et al.
Published: (2025)
Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
by: Wei, Wenda, et al.
Published: (2025)
by: Wei, Wenda, et al.
Published: (2025)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
by: Yang, Rushuai, et al.
Published: (2025)
by: Yang, Rushuai, et al.
Published: (2025)
SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning
by: Liu, Yihao, et al.
Published: (2025)
by: Liu, Yihao, et al.
Published: (2025)
Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents
by: Zhen, Shuai, et al.
Published: (2026)
by: Zhen, Shuai, et al.
Published: (2026)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
by: Guan, Zhong, et al.
Published: (2026)
by: Guan, Zhong, et al.
Published: (2026)
DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
by: Li, Yuanhao, et al.
Published: (2025)
by: Li, Yuanhao, et al.
Published: (2025)
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
by: Fang, Wenkai, et al.
Published: (2025)
by: Fang, Wenkai, et al.
Published: (2025)
RL2: Reinforce Large Language Model to Assist Safe Reinforcement Learning for Energy Management of Active Distribution Networks
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
by: Kim, Woosung, et al.
Published: (2025)
by: Kim, Woosung, et al.
Published: (2025)
Similar Items
-
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
by: Zhang, Xikai, et al.
Published: (2025) -
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026) -
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
by: Cui, Yitong, et al.
Published: (2025) -
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
by: Ren, Yanwei, et al.
Published: (2026) -
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
by: Zhang, Haotian, et al.
Published: (2025)