Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goldie, Anna, Mirhoseini, Azalia, Zhou, Hao, Cai, Irene, Manning, Christopher D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
von: Goldie, Anna, et al.
Veröffentlicht: (2024)
von: Goldie, Anna, et al.
Veröffentlicht: (2024)
On the Role of Temperature Sampling in Test-Time Scaling
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
von: Du, Weihua, et al.
Veröffentlicht: (2025)
von: Du, Weihua, et al.
Veröffentlicht: (2025)
Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2025)
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
von: Costello, Caia, et al.
Veröffentlicht: (2025)
von: Costello, Caia, et al.
Veröffentlicht: (2025)
SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
von: Biju, Emil, et al.
Veröffentlicht: (2025)
von: Biju, Emil, et al.
Veröffentlicht: (2025)
Cartridges: Lightweight and general-purpose long context representations via self-study
von: Eyuboglu, Sabri, et al.
Veröffentlicht: (2025)
von: Eyuboglu, Sabri, et al.
Veröffentlicht: (2025)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2025)
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2025)
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
von: Winston, Caleb, et al.
Veröffentlicht: (2026)
von: Winston, Caleb, et al.
Veröffentlicht: (2026)
Reasoning-Driven Synthetic Data Generation and Evaluation
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
von: Wang, Huaijie, et al.
Veröffentlicht: (2024)
von: Wang, Huaijie, et al.
Veröffentlicht: (2024)
Synthetic Data RL: Task Definition Is All You Need
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
Learning from Synthetic Data Improves Multi-hop Reasoning
von: Kabra, Anmol, et al.
Veröffentlicht: (2026)
von: Kabra, Anmol, et al.
Veröffentlicht: (2026)
Archon: An Architecture Search Framework for Inference-Time Techniques
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
ToolRL: Reward is All Tool Learning Needs
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
von: Ilin, Aleksei, et al.
Veröffentlicht: (2025)
von: Ilin, Aleksei, et al.
Veröffentlicht: (2025)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
OpenJarvis: Personal AI, On Personal Devices
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2026)
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2026)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
von: Kapoor, Vansh, et al.
Veröffentlicht: (2026)
von: Kapoor, Vansh, et al.
Veröffentlicht: (2026)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
von: Kulkarni, Anay, et al.
Veröffentlicht: (2026)
von: Kulkarni, Anay, et al.
Veröffentlicht: (2026)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
von: Zha, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zha, Kaiwen, et al.
Veröffentlicht: (2025)
CHESS: Contextual Harnessing for Efficient SQL Synthesis
von: Talaei, Shayan, et al.
Veröffentlicht: (2024)
von: Talaei, Shayan, et al.
Veröffentlicht: (2024)
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
von: Brown, Bradley, et al.
Veröffentlicht: (2024)
von: Brown, Bradley, et al.
Veröffentlicht: (2024)
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
von: Liu, Yuliang, et al.
Veröffentlicht: (2025)
von: Liu, Yuliang, et al.
Veröffentlicht: (2025)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
von: Wang, Dong, et al.
Veröffentlicht: (2025)
von: Wang, Dong, et al.
Veröffentlicht: (2025)
Multi-Step Reasoning with Large Language Models, a Survey
von: Plaat, Aske, et al.
Veröffentlicht: (2024)
von: Plaat, Aske, et al.
Veröffentlicht: (2024)
ProcBench: Benchmark for Multi-Step Reasoning and Following Procedure
von: Fujisawa, Ippei, et al.
Veröffentlicht: (2024)
von: Fujisawa, Ippei, et al.
Veröffentlicht: (2024)
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning
von: Shin, Kwan Soo
Veröffentlicht: (2026)
von: Shin, Kwan Soo
Veröffentlicht: (2026)
CALICO: Conversational Agent Localization via Synthetic Data Generation
von: Rosenbaum, Andy, et al.
Veröffentlicht: (2024)
von: Rosenbaum, Andy, et al.
Veröffentlicht: (2024)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
von: Park, Jinyoung, et al.
Veröffentlicht: (2023)
von: Park, Jinyoung, et al.
Veröffentlicht: (2023)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
von: Wang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Wang, Zhiwei, et al.
Veröffentlicht: (2024)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models
von: Yuan, Yefeng, et al.
Veröffentlicht: (2024)
von: Yuan, Yefeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
von: Goldie, Anna, et al.
Veröffentlicht: (2024) -
On the Role of Temperature Sampling in Test-Time Scaling
von: Wu, Yuheng, et al.
Veröffentlicht: (2025) -
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
von: Du, Weihua, et al.
Veröffentlicht: (2025) -
Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2025) -
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
von: Costello, Caia, et al.
Veröffentlicht: (2025)