A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Kaiwen, Liang, Dawen, Kallus, Nathan, Sun, Wen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Central Role of the Loss Function in Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
by: Lee, Jane H., et al.
Published: (2025)
by: Lee, Jane H., et al.
Published: (2025)
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2024)
by: Ayoub, Alex, et al.
Published: (2024)
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
by: Ghosh, Ayon, et al.
Published: (2024)
by: Ghosh, Ayon, et al.
Published: (2024)
On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents
by: Mortensen, Oliver, et al.
Published: (2026)
by: Mortensen, Oliver, et al.
Published: (2026)
Optimized Certainty Equivalent Risk-Controlling Prediction Sets
by: Huang, Jiayi, et al.
Published: (2026)
by: Huang, Jiayi, et al.
Published: (2026)
Does Weighting Improve Matrix Factorization for Recommender Systems?
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
by: Bennett, Andrew, et al.
Published: (2024)
by: Bennett, Andrew, et al.
Published: (2024)
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
by: Kallus, Nathan
Published: (2025)
by: Kallus, Nathan
Published: (2025)
Bellman Calibration for $V$-Learning in Offline Reinforcement Learning
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach
by: Hao, Guang-Yuan, et al.
Published: (2026)
by: Hao, Guang-Yuan, et al.
Published: (2026)
Robust and Agnostic Learning of Conditional Distributional Treatment Effects
by: Kallus, Nathan, et al.
Published: (2022)
by: Kallus, Nathan, et al.
Published: (2022)
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
by: McInerney, James, et al.
Published: (2024)
by: McInerney, James, et al.
Published: (2024)
Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Inverse Reinforcement Learning with Just Classification and a Few Regressions
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Estimating Heterogeneous Treatment Effects by Combining Weak Instruments and Observational Data
by: Oprescu, Miruna, et al.
Published: (2024)
by: Oprescu, Miruna, et al.
Published: (2024)
On the role of surrogates in the efficient estimation of treatment effects with limited outcome data
by: Kallus, Nathan, et al.
Published: (2020)
by: Kallus, Nathan, et al.
Published: (2020)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
by: Wang, Kaiwen, et al.
Published: (2025)
by: Wang, Kaiwen, et al.
Published: (2025)
Contextual Linear Optimization with Partial Feedback
by: Hu, Yichun, et al.
Published: (2024)
by: Hu, Yichun, et al.
Published: (2024)
Optimization of Epsilon-Greedy Exploration
by: Che, Ethan, et al.
Published: (2025)
by: Che, Ethan, et al.
Published: (2025)
A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
by: Kalyanakrishnan, Shivaram, et al.
Published: (2025)
by: Kalyanakrishnan, Shivaram, et al.
Published: (2025)
Semiparametric Double Reinforcement Learning with Applications to Long-Term Causal Inference
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Demistifying Inference after Adaptive Experiments
by: Bibaut, Aurélien, et al.
Published: (2024)
by: Bibaut, Aurélien, et al.
Published: (2024)
Multi-Armed Bandits with Interference
by: Jia, Su, et al.
Published: (2024)
by: Jia, Su, et al.
Published: (2024)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
by: Kausik, Chinmaya, et al.
Published: (2026)
by: Kausik, Chinmaya, et al.
Published: (2026)
Exploration in the Limit
by: Cho, Brian M., et al.
Published: (2025)
by: Cho, Brian M., et al.
Published: (2025)
Peeking with PEAK: Sequential, Nonparametric Composite Hypothesis Tests for Means of Multiple Data Streams
by: Cho, Brian, et al.
Published: (2024)
by: Cho, Brian, et al.
Published: (2024)
Is Risk-Sensitive Reinforcement Learning Properly Resolved?
by: Zhou, Ruiwen, et al.
Published: (2023)
by: Zhou, Ruiwen, et al.
Published: (2023)
Beyond Non-Degeneracy: Revisiting Certainty Equivalent Heuristic for Online Linear Programming
by: Chen, Yilun, et al.
Published: (2025)
by: Chen, Yilun, et al.
Published: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
by: Zhou, Jin Peng, et al.
Published: (2025)
by: Zhou, Jin Peng, et al.
Published: (2025)
Entropy After </Think> for reasoning model early exiting
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Is Cosine-Similarity of Embeddings Really About Similarity?
by: Steck, Harald, et al.
Published: (2024)
by: Steck, Harald, et al.
Published: (2024)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting
by: Moghimi, Mehrdad, et al.
Published: (2026)
by: Moghimi, Mehrdad, et al.
Published: (2026)
SNPL: Simultaneous Policy Learning and Evaluation for Safe Multi-Objective Policy Improvement
by: Cho, Brian, et al.
Published: (2025)
by: Cho, Brian, et al.
Published: (2025)
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Similar Items
-
The Central Role of the Loss Function in Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024) -
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
by: Lee, Jane H., et al.
Published: (2025) -
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024) -
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2024) -
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)