More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Kaiwen, Oertell, Owen, Agarwal, Alekh, Kallus, Nathan, Sun, Wen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Central Role of the Loss Function in Reinforcement Learning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
Convergence Of Consistency Model With Multistep Sampling Under General Data Assumptions
von: Chen, Yiding, et al.
Veröffentlicht: (2025)
von: Chen, Yiding, et al.
Veröffentlicht: (2025)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
von: Bennett, Andrew, et al.
Veröffentlicht: (2024)
von: Bennett, Andrew, et al.
Veröffentlicht: (2024)
Robust and Agnostic Learning of Conditional Distributional Treatment Effects
von: Kallus, Nathan, et al.
Veröffentlicht: (2022)
von: Kallus, Nathan, et al.
Veröffentlicht: (2022)
Bellman Calibration for $V$-Learning in Offline Reinforcement Learning
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
LLMs Can Learn to Reason Via Off-Policy RL
von: Ritter, Daniel, et al.
Veröffentlicht: (2026)
von: Ritter, Daniel, et al.
Veröffentlicht: (2026)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
von: Swamy, Gokul, et al.
Veröffentlicht: (2024)
von: Swamy, Gokul, et al.
Veröffentlicht: (2024)
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
von: Kallus, Nathan
Veröffentlicht: (2025)
von: Kallus, Nathan
Veröffentlicht: (2025)
Efficient Controllable Diffusion via Optimal Classifier Guidance
von: Oertell, Owen, et al.
Veröffentlicht: (2025)
von: Oertell, Owen, et al.
Veröffentlicht: (2025)
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
von: McInerney, James, et al.
Veröffentlicht: (2024)
von: McInerney, James, et al.
Veröffentlicht: (2024)
Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Model-based RL as a Minimalist Approach to Horizon-Free and Second-Order Bounds
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
Inverse Reinforcement Learning with Just Classification and a Few Regressions
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Offline Imitation Learning from Multiple Baselines with Applications to Compiler Optimization
von: Marinov, Teodor V., et al.
Veröffentlicht: (2024)
von: Marinov, Teodor V., et al.
Veröffentlicht: (2024)
Estimating Heterogeneous Treatment Effects by Combining Weak Instruments and Observational Data
von: Oprescu, Miruna, et al.
Veröffentlicht: (2024)
von: Oprescu, Miruna, et al.
Veröffentlicht: (2024)
On the role of surrogates in the efficient estimation of treatment effects with limited outcome data
von: Kallus, Nathan, et al.
Veröffentlicht: (2020)
von: Kallus, Nathan, et al.
Veröffentlicht: (2020)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach
von: Hao, Guang-Yuan, et al.
Veröffentlicht: (2026)
von: Hao, Guang-Yuan, et al.
Veröffentlicht: (2026)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
Preserving Expert-Level Privacy in Offline Reinforcement Learning
von: Sharma, Navodita, et al.
Veröffentlicht: (2024)
von: Sharma, Navodita, et al.
Veröffentlicht: (2024)
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
Semiparametric Double Reinforcement Learning with Applications to Long-Term Causal Inference
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Design Considerations in Offline Preference-based RL
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
Demistifying Inference after Adaptive Experiments
von: Bibaut, Aurélien, et al.
Veröffentlicht: (2024)
von: Bibaut, Aurélien, et al.
Veröffentlicht: (2024)
REBEL: Reinforcement Learning via Regressing Relative Rewards
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement Learning
von: Sun, Ke, et al.
Veröffentlicht: (2021)
von: Sun, Ke, et al.
Veröffentlicht: (2021)
Multi-Armed Bandits with Interference
von: Jia, Su, et al.
Veröffentlicht: (2024)
von: Jia, Su, et al.
Veröffentlicht: (2024)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
TurboHopp: Accelerated Molecule Scaffold Hopping with Consistency Models
von: Yoo, Kiwoong, et al.
Veröffentlicht: (2024)
von: Yoo, Kiwoong, et al.
Veröffentlicht: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Mitigating Preference Hacking in Policy Optimization with Pessimism
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
Exploration in the Limit
von: Cho, Brian M., et al.
Veröffentlicht: (2025)
von: Cho, Brian M., et al.
Veröffentlicht: (2025)
Peeking with PEAK: Sequential, Nonparametric Composite Hypothesis Tests for Means of Multiple Data Streams
von: Cho, Brian, et al.
Veröffentlicht: (2024)
von: Cho, Brian, et al.
Veröffentlicht: (2024)
Entropy After </Think> for reasoning model early exiting
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Central Role of the Loss Function in Reinforcement Learning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024) -
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024) -
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022) -
Convergence Of Consistency Model With Multistep Sampling Under General Data Assumptions
von: Chen, Yiding, et al.
Veröffentlicht: (2025) -
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
von: Bennett, Andrew, et al.
Veröffentlicht: (2024)