From Robotics to Sepsis Treatment: Offline RL via Geometric Pessimism
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Wanjari, Sarthak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Tale of Two Cities: Pessimism and Opportunism in Offline Dynamic Pricing
von: Bian, Zeyu, et al.
Veröffentlicht: (2024)
von: Bian, Zeyu, et al.
Veröffentlicht: (2024)
Beyond Pessimism: Offline Learning in KL-regularized Games
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
von: Roknilamouki, Amirhossein, et al.
Veröffentlicht: (2026)
von: Roknilamouki, Amirhossein, et al.
Veröffentlicht: (2026)
Pessimism-Free Offline Learning in General-Sum Games via KL Regularization
von: Chen, Claire, et al.
Veröffentlicht: (2026)
von: Chen, Claire, et al.
Veröffentlicht: (2026)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
von: Zhang, Dake, et al.
Veröffentlicht: (2024)
von: Zhang, Dake, et al.
Veröffentlicht: (2024)
Attention-Based Offline Reinforcement Learning and Clustering for Interpretable Sepsis Treatment
von: Kumar, Punit, et al.
Veröffentlicht: (2026)
von: Kumar, Punit, et al.
Veröffentlicht: (2026)
Exploring a Graph-based Approach to Offline Reinforcement Learning for Sepsis Treatment
von: Khakharova, Taisiya, et al.
Veröffentlicht: (2025)
von: Khakharova, Taisiya, et al.
Veröffentlicht: (2025)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
von: Aouali, Imad, et al.
Veröffentlicht: (2024)
von: Aouali, Imad, et al.
Veröffentlicht: (2024)
Stable CDE Autoencoders with Acuity Regularization for Offline Reinforcement Learning in Sepsis Treatment
von: Gao, Yue
Veröffentlicht: (2025)
von: Gao, Yue
Veröffentlicht: (2025)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
von: Yu, Zhuohao, et al.
Veröffentlicht: (2026)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2026)
The Virtues of Pessimism in Inverse Reinforcement Learning
von: Wu, David, et al.
Veröffentlicht: (2024)
von: Wu, David, et al.
Veröffentlicht: (2024)
Language-Conditioned Offline RL for Multi-Robot Navigation
von: Morad, Steven, et al.
Veröffentlicht: (2024)
von: Morad, Steven, et al.
Veröffentlicht: (2024)
BoxRL-NNV: Boxed Refinement of Latin Hypercube Samples for Neural Network Verification
von: Das, Sarthak
Veröffentlicht: (2025)
von: Das, Sarthak
Veröffentlicht: (2025)
Offline RL via Feature-Occupancy Gradient Ascent
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
von: Lu, Miao, et al.
Veröffentlicht: (2022)
von: Lu, Miao, et al.
Veröffentlicht: (2022)
Agentic Planning with Reasoning for Image Styling via Offline RL
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2026)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2026)
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning
von: Kobayashi, Taisuke
Veröffentlicht: (2024)
von: Kobayashi, Taisuke
Veröffentlicht: (2024)
Improving Offline RL by Blending Heuristics
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
OMG-RL:Offline Model-based Guided Reward Learning for Heparin Treatment
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
From Path Signatures to Sequential Modeling: Incremental Signature Contributions for Offline RL
von: Zhao, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziyi, et al.
Veröffentlicht: (2026)
Mitigating Preference Hacking in Policy Optimization with Pessimism
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data
von: Zheng, Chongyi, et al.
Veröffentlicht: (2023)
von: Zheng, Chongyi, et al.
Veröffentlicht: (2023)
When Are RL Hyperparameters Benign? A Study in Offline Goal-Conditioned RL
von: Töpperwien, Jan Malte, et al.
Veröffentlicht: (2026)
von: Töpperwien, Jan Malte, et al.
Veröffentlicht: (2026)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
von: Beck, Jacob
Veröffentlicht: (2025)
von: Beck, Jacob
Veröffentlicht: (2025)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
von: Wang, Qi, et al.
Veröffentlicht: (2023)
von: Wang, Qi, et al.
Veröffentlicht: (2023)
Selective Uncertainty Propagation in Offline RL
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
Decoupled Prioritized Resampling for Offline RL
von: Yue, Yang, et al.
Veröffentlicht: (2023)
von: Yue, Yang, et al.
Veröffentlicht: (2023)
Augmenting Offline RL with Unlabeled Data
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets
von: Gupta, Aaryan, et al.
Veröffentlicht: (2025)
von: Gupta, Aaryan, et al.
Veröffentlicht: (2025)
Reinformer: Max-Return Sequence Modeling for Offline RL
von: Zhuang, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhuang, Zifeng, et al.
Veröffentlicht: (2024)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
Action-Free Offline-to-Online RL via Discretised State Policies
von: Neggatu, Natinael Solomon, et al.
Veröffentlicht: (2026)
von: Neggatu, Natinael Solomon, et al.
Veröffentlicht: (2026)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
von: Hsu, Hsiang, et al.
Veröffentlicht: (2026)
von: Hsu, Hsiang, et al.
Veröffentlicht: (2026)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
A Tractable Inference Perspective of Offline RL
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
Design Considerations in Offline Preference-based RL
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
Are Expressive Models Truly Necessary for Offline RL?
von: Wang, Guan, et al.
Veröffentlicht: (2024)
von: Wang, Guan, et al.
Veröffentlicht: (2024)
OGBench: Benchmarking Offline Goal-Conditioned RL
von: Park, Seohong, et al.
Veröffentlicht: (2024)
von: Park, Seohong, et al.
Veröffentlicht: (2024)
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
von: Yu, Kihyun, et al.
Veröffentlicht: (2024)
von: Yu, Kihyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Tale of Two Cities: Pessimism and Opportunism in Offline Dynamic Pricing
von: Bian, Zeyu, et al.
Veröffentlicht: (2024) -
Beyond Pessimism: Offline Learning in KL-regularized Games
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026) -
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
von: Roknilamouki, Amirhossein, et al.
Veröffentlicht: (2026) -
Pessimism-Free Offline Learning in General-Sum Games via KL Regularization
von: Chen, Claire, et al.
Veröffentlicht: (2026) -
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
von: Zhang, Dake, et al.
Veröffentlicht: (2024)