Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qu, Yuxiao, Yang, Matthew Y. R., Setlur, Amrith, Tunstall, Lewis, Beeching, Edward Emanuel, Salakhutdinov, Ruslan, Kumar, Aviral |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026)
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
von: LM-Provers, et al.
Veröffentlicht: (2026)
von: LM-Provers, et al.
Veröffentlicht: (2026)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
von: Wu, Ian, et al.
Veröffentlicht: (2026)
von: Wu, Ian, et al.
Veröffentlicht: (2026)
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
von: Yang, Matthew Y. R., et al.
Veröffentlicht: (2026)
von: Yang, Matthew Y. R., et al.
Veröffentlicht: (2026)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
von: Cheng, Zhoujun, et al.
Veröffentlicht: (2026)
von: Cheng, Zhoujun, et al.
Veröffentlicht: (2026)
Multi-Agent Computer Use
von: Koh, Jing Yu, et al.
Veröffentlicht: (2026)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2026)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
von: Niklaus, Joel, et al.
Veröffentlicht: (2026)
von: Niklaus, Joel, et al.
Veröffentlicht: (2026)
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
CaRT: Teaching LLM Agents to Know When They Know Enough
von: Liu, Grace, et al.
Veröffentlicht: (2025)
von: Liu, Grace, et al.
Veröffentlicht: (2025)
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
von: Setlur, Amrith, et al.
Veröffentlicht: (2026)
von: Setlur, Amrith, et al.
Veröffentlicht: (2026)
Scaling Test-Time Compute for Agentic Coding
von: Kim, Joongwon, et al.
Veröffentlicht: (2026)
von: Kim, Joongwon, et al.
Veröffentlicht: (2026)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?
von: Kang, Katie, et al.
Veröffentlicht: (2024)
von: Kang, Katie, et al.
Veröffentlicht: (2024)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
von: Jang, Lawrence Keunho, et al.
Veröffentlicht: (2026)
von: Jang, Lawrence Keunho, et al.
Veröffentlicht: (2026)
Automatic Question-Answer Generation for Long-Tail Knowledge
von: Kumar, Rohan, et al.
Veröffentlicht: (2024)
von: Kumar, Rohan, et al.
Veröffentlicht: (2024)
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
von: Lu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Lu, Yuxiao, et al.
Veröffentlicht: (2024)
Lower Bounds for Public-Private Learning under Distribution Shift
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
Tree Search for Language Model Agents
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
von: Li, Yafu, et al.
Veröffentlicht: (2025)
von: Li, Yafu, et al.
Veröffentlicht: (2025)
Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning
von: Singh, Navan Preet, et al.
Veröffentlicht: (2026)
von: Singh, Navan Preet, et al.
Veröffentlicht: (2026)
Parameter-Efficient Fine-Tuning for Foundation Models
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
In-Context Fine-Tuning for Time-Series Foundation Models
von: Das, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Das, Abhimanyu, et al.
Veröffentlicht: (2024)
ReFT: Reasoning with Reinforced Fine-Tuning
von: Luong, Trung Quoc, et al.
Veröffentlicht: (2024)
von: Luong, Trung Quoc, et al.
Veröffentlicht: (2024)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
von: Yu, Haofei, et al.
Veröffentlicht: (2023)
von: Yu, Haofei, et al.
Veröffentlicht: (2023)
Prior Prompt Engineering for Reinforcement Fine-Tuning
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2025)
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2025)
UFT: Unifying Supervised and Reinforcement Fine-Tuning
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
Self-Evolution Fine-Tuning for Policy Optimization
von: Chen, Ruijun, et al.
Veröffentlicht: (2024)
von: Chen, Ruijun, et al.
Veröffentlicht: (2024)
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
von: Khadangi, Afshin, et al.
Veröffentlicht: (2025)
von: Khadangi, Afshin, et al.
Veröffentlicht: (2025)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
Benchmark Test-Time Scaling of General LLM Agents
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
von: Wu, Mian, et al.
Veröffentlicht: (2025)
von: Wu, Mian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026) -
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
von: LM-Provers, et al.
Veröffentlicht: (2026) -
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025) -
Scaling Test-Time Compute Without Verification or RL is Suboptimal
von: Setlur, Amrith, et al.
Veröffentlicht: (2025) -
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
von: Wu, Ian, et al.
Veröffentlicht: (2026)