Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
Fuente:
arXiv
Saved in:
| Main Authors: | Setlur, Amrith, Wang, Zijian, Cohen, Andrew, Rashidinejad, Paria, Xie, Sang Michael |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
by: Qu, Yuxiao, et al.
Published: (2026)
by: Qu, Yuxiao, et al.
Published: (2026)
Solve the Loop: Attractor Models for Language and Reasoning
by: Fein-Ashley, Jacob, et al.
Published: (2026)
by: Fein-Ashley, Jacob, et al.
Published: (2026)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
by: Rashidinejad, Paria, et al.
Published: (2024)
by: Rashidinejad, Paria, et al.
Published: (2024)
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
by: LM-Provers, et al.
Published: (2026)
by: LM-Provers, et al.
Published: (2026)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
by: Kim, Minwu, et al.
Published: (2026)
by: Kim, Minwu, et al.
Published: (2026)
Efficiency-Effectiveness Reranking FLOPs for LLM-based Rerankers
by: Peng, Zhiyuan, et al.
Published: (2025)
by: Peng, Zhiyuan, et al.
Published: (2025)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
by: Yang, Matthew Y. R., et al.
Published: (2026)
by: Yang, Matthew Y. R., et al.
Published: (2026)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
by: Setlur, Amrith, et al.
Published: (2024)
by: Setlur, Amrith, et al.
Published: (2024)
CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing
by: Ikram, Zarif, et al.
Published: (2026)
by: Ikram, Zarif, et al.
Published: (2026)
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
by: Pouplin, Thomas, et al.
Published: (2024)
by: Pouplin, Thomas, et al.
Published: (2024)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
by: Cheng, Zhoujun, et al.
Published: (2026)
by: Cheng, Zhoujun, et al.
Published: (2026)
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
by: Cen, Zhepeng, et al.
Published: (2025)
by: Cen, Zhepeng, et al.
Published: (2025)
Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
by: Portes, Jacob, et al.
Published: (2025)
by: Portes, Jacob, et al.
Published: (2025)
Path-Consistency with Prefix Enhancement for Efficient Inference in LLMs
by: Zhu, Jiace, et al.
Published: (2024)
by: Zhu, Jiace, et al.
Published: (2024)
Detecting Prefix Bias in LLM-based Reward Models
by: Kumar, Ashwin, et al.
Published: (2025)
by: Kumar, Ashwin, et al.
Published: (2025)
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
by: Harary, Sapir, et al.
Published: (2025)
by: Harary, Sapir, et al.
Published: (2025)
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
by: Wang, Yuyao, et al.
Published: (2025)
by: Wang, Yuyao, et al.
Published: (2025)
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
by: Guan, Xin, et al.
Published: (2026)
by: Guan, Xin, et al.
Published: (2026)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
by: Gao, Jiaxuan, et al.
Published: (2025)
by: Gao, Jiaxuan, et al.
Published: (2025)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
Solving an Open Problem in Theoretical Physics using AI-Assisted Discovery
by: Brenner, Michael P., et al.
Published: (2026)
by: Brenner, Michael P., et al.
Published: (2026)
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
by: Liu, Jiacheng, et al.
Published: (2023)
by: Liu, Jiacheng, et al.
Published: (2023)
Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL
by: Choi, Yunseon, et al.
Published: (2024)
by: Choi, Yunseon, et al.
Published: (2024)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
by: Lin, Xiaofeng, et al.
Published: (2026)
by: Lin, Xiaofeng, et al.
Published: (2026)
Editing Knowledge Representation of Language Model via Rephrased Prefix Prompts
by: Cai, Yuchen, et al.
Published: (2024)
by: Cai, Yuchen, et al.
Published: (2024)
The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus
by: Jindal, Ishan, et al.
Published: (2026)
by: Jindal, Ishan, et al.
Published: (2026)
Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
by: Cai, Min, et al.
Published: (2024)
by: Cai, Min, et al.
Published: (2024)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
by: Zhang, Jiazheng, et al.
Published: (2026)
by: Zhang, Jiazheng, et al.
Published: (2026)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
Similar Items
-
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
by: Qu, Yuxiao, et al.
Published: (2026) -
Solve the Loop: Attractor Models for Language and Reasoning
by: Fein-Ashley, Jacob, et al.
Published: (2026) -
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025) -
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
by: Rashidinejad, Paria, et al.
Published: (2024) -
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
by: LM-Provers, et al.
Published: (2026)