Saved in:
| Main Authors: | Liu, Zhihan, Guan, Lin, Nie, Yixin, Zhang, Kai, Hao, Zhuoqun, Chen, Lin, Celikyilmaz, Asli, Wang, Zhaoran, Zhang, Na |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.18217 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Majority is not always right: RL training for solution aggregation
by: Zhao, Wenting, et al.
Published: (2025)
by: Zhao, Wenting, et al.
Published: (2025)
Open-Domain Text Evaluation via Contrastive Distribution Methods
by: Lu, Sidi, et al.
Published: (2023)
by: Lu, Sidi, et al.
Published: (2023)
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
by: Liu, Zhihan, et al.
Published: (2023)
by: Liu, Zhihan, et al.
Published: (2023)
How Can LLM Guide RL? A Value-Based Approach
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
by: Saha, Swarnadeep, et al.
Published: (2023)
by: Saha, Swarnadeep, et al.
Published: (2023)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
by: Li, Shuyue Stella, et al.
Published: (2026)
by: Li, Shuyue Stella, et al.
Published: (2026)
Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided Refinement
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
Can RL Improve Generalization of LLM Agents? An Empirical Study
by: Xi, Zhiheng, et al.
Published: (2026)
by: Xi, Zhiheng, et al.
Published: (2026)
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
by: Liu, Jiacheng, et al.
Published: (2023)
by: Liu, Jiacheng, et al.
Published: (2023)
Learning a Game by Paying the Agents
by: Zhang, Brian Hu, et al.
Published: (2025)
by: Zhang, Brian Hu, et al.
Published: (2025)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval
by: Lin, Hao, et al.
Published: (2025)
by: Lin, Hao, et al.
Published: (2025)
TaxDiff: Taxonomic-Guided Diffusion Model for Protein Sequence Generation
by: Zongying, Lin, et al.
Published: (2024)
by: Zongying, Lin, et al.
Published: (2024)
The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
by: Yang, Kevin, et al.
Published: (2023)
by: Yang, Kevin, et al.
Published: (2023)
RL-LLM-DT: An Automatic Decision Tree Generation Method Based on RL Evaluation and LLM Enhancement
by: Lin, Junjie, et al.
Published: (2024)
by: Lin, Junjie, et al.
Published: (2024)
Paying Alignment Tax with Contrastive Learning
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
The Autonomy Tax: Defense Training Breaks LLM Agents
by: Li, Shawn, et al.
Published: (2026)
by: Li, Shawn, et al.
Published: (2026)
Less Data Less Tokens: Multilingual Unification Learning for Efficient Test-Time Reasoning in LLMs
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
Cold-Start Personalization via Training-Free Priors from Structured World Models
by: Bose, Avinandan, et al.
Published: (2026)
by: Bose, Avinandan, et al.
Published: (2026)
How Useful Is Cross-Domain Generalization for Training LLM Monitors?
by: Martin, Sam, et al.
Published: (2026)
by: Martin, Sam, et al.
Published: (2026)
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents
by: Zhang, Kaituo, et al.
Published: (2026)
by: Zhang, Kaituo, et al.
Published: (2026)
Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following
by: Yang, Yuxiao, et al.
Published: (2024)
by: Yang, Yuxiao, et al.
Published: (2024)
Pay Attention and Move Better: Harnessing Attention for Interactive Motion Generation and Training-free Editing
by: Chen, Ling-Hao, et al.
Published: (2024)
by: Chen, Ling-Hao, et al.
Published: (2024)
Sobolev-Regularised Flow Matching for Temporally Smooth Robot Trajectory Generation
by: Zhang, Zhihan
Published: (2026)
by: Zhang, Zhihan
Published: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding
by: Yang, Xu, et al.
Published: (2026)
by: Yang, Xu, et al.
Published: (2026)
reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs
by: Wu, Zhaofeng, et al.
Published: (2025)
by: Wu, Zhaofeng, et al.
Published: (2025)
Learning to Interrupt in Language-based Multi-agent Communication
by: Wang, Danqing, et al.
Published: (2026)
by: Wang, Danqing, et al.
Published: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Training microrobots to swim by a large language model
by: Xu, Zhuoqun, et al.
Published: (2024)
by: Xu, Zhuoqun, et al.
Published: (2024)
Simulating Macroeconomic Expectations in Survey Experiments with LLM-based Economic Agents
by: Lin, Jianhao, et al.
Published: (2025)
by: Lin, Jianhao, et al.
Published: (2025)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
by: Li, Pengbo, et al.
Published: (2026)
by: Li, Pengbo, et al.
Published: (2026)
Similar Items
-
The Majority is not always right: RL training for solution aggregation
by: Zhao, Wenting, et al.
Published: (2025) -
Open-Domain Text Evaluation via Contrastive Distribution Methods
by: Lu, Sidi, et al.
Published: (2023) -
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
by: Liu, Zhihan, et al.
Published: (2023) -
How Can LLM Guide RL? A Value-Based Approach
by: Zhang, Shenao, et al.
Published: (2024) -
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024)