Multi-Turn Code Generation Through Single-Step Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Arnav Kumar, Gonzalez-Pumariega, Gonzalo, Chen, Wayne, Rush, Alexander M, Zhao, Wenting, Choudhury, Sanjiban |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Query-Efficient Planning with Language Models
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2024)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2024)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Process Reward Models for LLM Agents: Practical Framework and Directions
by: Choudhury, Sanjiban
Published: (2025)
by: Choudhury, Sanjiban
Published: (2025)
Great Memory, Shallow Reasoning: Limits of $k$NN-LMs
by: Geng, Shangyi, et al.
Published: (2024)
by: Geng, Shangyi, et al.
Published: (2024)
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Non-Adversarial Inverse Reinforcement Learning via Successor Feature Matching
by: Jain, Arnav Kumar, et al.
Published: (2024)
by: Jain, Arnav Kumar, et al.
Published: (2024)
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
by: Wan, Yanming, et al.
Published: (2025)
by: Wan, Yanming, et al.
Published: (2025)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
by: Chen, Xingwu, et al.
Published: (2026)
by: Chen, Xingwu, et al.
Published: (2026)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
by: Li, Yubo, et al.
Published: (2025)
by: Li, Yubo, et al.
Published: (2025)
Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback
by: Choudhury, Sanjiban, et al.
Published: (2024)
by: Choudhury, Sanjiban, et al.
Published: (2024)
CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code
by: Ahmed, Tasnim, et al.
Published: (2025)
by: Ahmed, Tasnim, et al.
Published: (2025)
Integrating Planning into Single-Turn Long-Form Text Generation
by: Liang, Yi, et al.
Published: (2024)
by: Liang, Yi, et al.
Published: (2024)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
by: Zhang, Ruiyi, et al.
Published: (2025)
by: Zhang, Ruiyi, et al.
Published: (2025)
Large Language Models for Single-Step and Multi-Step Flight Trajectory Prediction
by: Luo, Kaiwei, et al.
Published: (2025)
by: Luo, Kaiwei, et al.
Published: (2025)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
by: Zhong, Haitian, et al.
Published: (2026)
by: Zhong, Haitian, et al.
Published: (2026)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
by: Chen, Ziru, et al.
Published: (2026)
by: Chen, Ziru, et al.
Published: (2026)
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
by: Lee, Young-Suk, et al.
Published: (2024)
by: Lee, Young-Suk, et al.
Published: (2024)
Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models
by: Wang, Teng, et al.
Published: (2025)
by: Wang, Teng, et al.
Published: (2025)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
by: Graf, Victoria, et al.
Published: (2026)
by: Graf, Victoria, et al.
Published: (2026)
Contextual Document Embeddings
by: Morris, John X., et al.
Published: (2024)
by: Morris, John X., et al.
Published: (2024)
Speak or Stay Silent: Context-Aware Turn-Taking in Multi-Party Dialogue
by: Bhagtani, Kratika, et al.
Published: (2026)
by: Bhagtani, Kratika, et al.
Published: (2026)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
by: Chen, Yi-Chun, et al.
Published: (2025)
by: Chen, Yi-Chun, et al.
Published: (2025)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
by: Kulkarni, Anay, et al.
Published: (2026)
by: Kulkarni, Anay, et al.
Published: (2026)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
by: Fan, Lishui, et al.
Published: (2025)
by: Fan, Lishui, et al.
Published: (2025)
SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering
by: Shi, Qiming, et al.
Published: (2026)
by: Shi, Qiming, et al.
Published: (2026)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
by: Chakraborty, Amartya, et al.
Published: (2025)
by: Chakraborty, Amartya, et al.
Published: (2025)
CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules
by: Le, Hung, et al.
Published: (2023)
by: Le, Hung, et al.
Published: (2023)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
by: Wu, Yuning, et al.
Published: (2026)
by: Wu, Yuning, et al.
Published: (2026)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
A Static and Dynamic Attention Framework for Multi Turn Dialogue Generation
by: Zhang, Wei-Nan, et al.
Published: (2024)
by: Zhang, Wei-Nan, et al.
Published: (2024)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
by: Goru, Ritesh, et al.
Published: (2025)
by: Goru, Ritesh, et al.
Published: (2025)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
by: Thornton, Scott
Published: (2025)
by: Thornton, Scott
Published: (2025)
MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
by: Mandal, Aishik, et al.
Published: (2025)
by: Mandal, Aishik, et al.
Published: (2025)
CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
by: Ding, Yifeng, et al.
Published: (2025)
by: Ding, Yifeng, et al.
Published: (2025)
Attributing Culture-Conditioned Generations to Pretraining Corpora
by: Li, Huihan, et al.
Published: (2024)
by: Li, Huihan, et al.
Published: (2024)
GRAM: A Generative Foundation Reward Model for Reward Generalization
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
Similar Items
-
Query-Efficient Planning with Language Models
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2024) -
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025) -
Process Reward Models for LLM Agents: Practical Framework and Directions
by: Choudhury, Sanjiban
Published: (2025) -
Great Memory, Shallow Reasoning: Limits of $k$NN-LMs
by: Geng, Shangyi, et al.
Published: (2024) -
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024)