TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Peiji, Li, Linyang, Sun, Handa, Mai, Wenjin, Chen, Yongkang, Li, Xiaozhe, Shen, Yue, Ma, Yichuan, Sun, Yiliu, Cao, Jiaxi, He, Zhishu, Wang, Bo, Zheng, Xiaoqing, Bi, Zhaori, Qiu, Xipeng, Guo, Qipeng, Chen, Kai, Lin, Dahua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
di: Ma, Yichuan, et al.
Pubblicazione: (2025)
di: Ma, Yichuan, et al.
Pubblicazione: (2025)
FastMCTS: A Simple Sampling Strategy for Data Synthesis
di: Li, Peiji, et al.
Pubblicazione: (2025)
di: Li, Peiji, et al.
Pubblicazione: (2025)
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
di: Li, Peiji, et al.
Pubblicazione: (2025)
di: Li, Peiji, et al.
Pubblicazione: (2025)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Case2Code: Scalable Synthetic Data for Code Generation
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
di: Sun, Yu, et al.
Pubblicazione: (2024)
di: Sun, Yu, et al.
Pubblicazione: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025)
di: Wang, Bo, et al.
Pubblicazione: (2025)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
Balanced Data Sampling for Language Model Training with Clustering
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
Importance of sample size determination for randomized controlled clinical trials for coronavirus disease 2019 antiviral therapies
di: Getu Zhaori
Pubblicazione: (2024)
di: Getu Zhaori
Pubblicazione: (2024)
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
di: Zhong, Yicheng, et al.
Pubblicazione: (2025)
di: Zhong, Yicheng, et al.
Pubblicazione: (2025)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration
di: Deng, Liyuan, et al.
Pubblicazione: (2026)
di: Deng, Liyuan, et al.
Pubblicazione: (2026)
COSMO-Agent: Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration
di: Deng, Liyuan, et al.
Pubblicazione: (2026)
di: Deng, Liyuan, et al.
Pubblicazione: (2026)
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
di: Li, Junbo, et al.
Pubblicazione: (2025)
di: Li, Junbo, et al.
Pubblicazione: (2025)
Phonon angular momentum induced by Terahertz electric field
di: Sun, Hong, et al.
Pubblicazione: (2025)
di: Sun, Hong, et al.
Pubblicazione: (2025)
GSC 08227-00723: An Unusually Large PSH Excess AH Pic Candidate
di: Li, Xin, et al.
Pubblicazione: (2026)
di: Li, Xin, et al.
Pubblicazione: (2026)
Identifying Semantic Induction Heads to Understand In-Context Learning
di: Ren, Jie, et al.
Pubblicazione: (2024)
di: Ren, Jie, et al.
Pubblicazione: (2024)
Data-free Weight Compress and Denoise for Large Language Models
di: Peng, Runyu, et al.
Pubblicazione: (2024)
di: Peng, Runyu, et al.
Pubblicazione: (2024)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
di: Lv, Kai, et al.
Pubblicazione: (2024)
di: Lv, Kai, et al.
Pubblicazione: (2024)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
di: Peng, Runyu, et al.
Pubblicazione: (2026)
di: Peng, Runyu, et al.
Pubblicazione: (2026)
MedKGI: Iterative Differential Diagnosis with Medical Knowledge Graphs and Information-Guided Inquiring
di: Wang, Qipeng, et al.
Pubblicazione: (2025)
di: Wang, Qipeng, et al.
Pubblicazione: (2025)
Can AI Assistants Know What They Don't Know?
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
di: Chen, Ziru, et al.
Pubblicazione: (2026)
di: Chen, Ziru, et al.
Pubblicazione: (2026)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2026)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2026)
Targeting protein condensation in cGAS‐STING signaling pathway
di: Yajie Li, et al.
Pubblicazione: (2024)
di: Yajie Li, et al.
Pubblicazione: (2024)
CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate
di: Sun, Yiliu, et al.
Pubblicazione: (2025)
di: Sun, Yiliu, et al.
Pubblicazione: (2025)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
di: Li, Aiden Yiliu, et al.
Pubblicazione: (2025)
di: Li, Aiden Yiliu, et al.
Pubblicazione: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
di: Liu, Jie, et al.
Pubblicazione: (2025)
di: Liu, Jie, et al.
Pubblicazione: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
di: Ding, Zheng, et al.
Pubblicazione: (2025)
di: Ding, Zheng, et al.
Pubblicazione: (2025)
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
Unified Active Retrieval for Retrieval Augmented Generation
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
di: Tong, Chengzhuo, et al.
Pubblicazione: (2025)
di: Tong, Chengzhuo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
di: Ma, Yichuan, et al.
Pubblicazione: (2026) -
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
di: Li, Xiaozhe, et al.
Pubblicazione: (2026) -
Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
di: Ma, Yichuan, et al.
Pubblicazione: (2026) -
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
di: Ma, Yichuan, et al.
Pubblicazione: (2025) -
FastMCTS: A Simple Sampling Strategy for Data Synthesis
di: Li, Peiji, et al.
Pubblicazione: (2025)