$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Jin Peng, Wang, Kaiwen, Chang, Jonathan, Gao, Zhaolin, Kallus, Nathan, Weinberger, Kilian Q., Brantley, Kianté, Sun, Wen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Value-Guided Search for Efficient Chain-of-Thought Reasoning
by: Wang, Kaiwen, et al.
Published: (2025)
by: Wang, Kaiwen, et al.
Published: (2025)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
by: Brantley, Kianté, et al.
Published: (2025)
by: Brantley, Kianté, et al.
Published: (2025)
Reviewer2: Optimizing Review Generation Through Prompt Generation
by: Gao, Zhaolin, et al.
Published: (2024)
by: Gao, Zhaolin, et al.
Published: (2024)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
by: Oertell, Owen, et al.
Published: (2024)
by: Oertell, Owen, et al.
Published: (2024)
LLMs Can Learn to Reason Via Off-Policy RL
by: Ritter, Daniel, et al.
Published: (2026)
by: Ritter, Daniel, et al.
Published: (2026)
$p1$: Better Prompt Optimization with Fewer Prompts
by: Gao, Zhaolin, et al.
Published: (2026)
by: Gao, Zhaolin, et al.
Published: (2026)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
by: Gao, Zhaolin, et al.
Published: (2024)
by: Gao, Zhaolin, et al.
Published: (2024)
Policy-Gradient Training of Language Models for Ranking
by: Gao, Ge, et al.
Published: (2023)
by: Gao, Ge, et al.
Published: (2023)
Expressive Value Learning for Scalable Offline Reinforcement Learning
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Adversarial Imitation Learning via Boosting
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
The Central Role of the Loss Function in Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
by: Wu, Anne, et al.
Published: (2024)
by: Wu, Anne, et al.
Published: (2024)
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Diffusing States and Matching Scores: A New Framework for Imitation Learning
by: Wu, Runzhe, et al.
Published: (2024)
by: Wu, Runzhe, et al.
Published: (2024)
Scaling Offline RL via Efficient and Expressive Shortcut Models
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
REBEL: Reinforcement Learning via Regressing Relative Rewards
by: Gao, Zhaolin, et al.
Published: (2024)
by: Gao, Zhaolin, et al.
Published: (2024)
Dataset Reset Policy Optimization for RLHF
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
Prompt Curriculum Learning for Efficient LLM Post-Training
by: Gao, Zhaolin, et al.
Published: (2025)
by: Gao, Zhaolin, et al.
Published: (2025)
Ranking with Long-Term Constraints
by: Brantley, Kianté, et al.
Published: (2023)
by: Brantley, Kianté, et al.
Published: (2023)
LLMs Are In-Context Bandit Reinforcement Learners
by: Monea, Giovanni, et al.
Published: (2024)
by: Monea, Giovanni, et al.
Published: (2024)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
by: Bennett, Andrew, et al.
Published: (2024)
by: Bennett, Andrew, et al.
Published: (2024)
Robust and Agnostic Learning of Conditional Distributional Treatment Effects
by: Kallus, Nathan, et al.
Published: (2022)
by: Kallus, Nathan, et al.
Published: (2022)
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
by: Kallus, Nathan
Published: (2025)
by: Kallus, Nathan
Published: (2025)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
by: Monea, Giovanni, et al.
Published: (2025)
by: Monea, Giovanni, et al.
Published: (2025)
Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
by: Wang, Qizhou, et al.
Published: (2025)
by: Wang, Qizhou, et al.
Published: (2025)
Orchestrating LLMs with Different Personalizations
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
INPROVF: Leveraging Large Language Models to Repair High-level Robot Controllers from Assumption Violations
by: Meng, Qian, et al.
Published: (2025)
by: Meng, Qian, et al.
Published: (2025)
Prescriptive Scaling Laws for Data Constrained Training
by: Lovelace, Justin, et al.
Published: (2026)
by: Lovelace, Justin, et al.
Published: (2026)
Near-Optimal Non-Parametric Sequential Tests and Confidence Sequences with Possibly Dependent Observations
by: Bibaut, Aurelien, et al.
Published: (2022)
by: Bibaut, Aurelien, et al.
Published: (2022)
On Speeding Up Language Model Evaluation
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
by: Wang, Zhixin, et al.
Published: (2025)
by: Wang, Zhixin, et al.
Published: (2025)
Demistifying Inference after Adaptive Experiments
by: Bibaut, Aurélien, et al.
Published: (2024)
by: Bibaut, Aurélien, et al.
Published: (2024)
Anytime-Valid Continuous-Time Confidence Processes for Inhomogeneous Poisson Processes
by: Lindon, Michael, et al.
Published: (2024)
by: Lindon, Michael, et al.
Published: (2024)
Estimating Heterogeneous Treatment Effects by Combining Weak Instruments and Observational Data
by: Oprescu, Miruna, et al.
Published: (2024)
by: Oprescu, Miruna, et al.
Published: (2024)
On the role of surrogates in the efficient estimation of treatment effects with limited outcome data
by: Kallus, Nathan, et al.
Published: (2020)
by: Kallus, Nathan, et al.
Published: (2020)
Similar Items
-
Value-Guided Search for Efficient Chain-of-Thought Reasoning
by: Wang, Kaiwen, et al.
Published: (2025) -
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
by: Brantley, Kianté, et al.
Published: (2025) -
Reviewer2: Optimizing Review Generation Through Prompt Generation
by: Gao, Zhaolin, et al.
Published: (2024) -
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
by: Oertell, Owen, et al.
Published: (2024) -
LLMs Can Learn to Reason Via Off-Policy RL
by: Ritter, Daniel, et al.
Published: (2026)