Saved in:
| Main Authors: | Deb, Rohan, Wright, Stephen J., Banerjee, Arindam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.22430 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Plan Before You Trade: Inference-Time Optimization for RL Trading Agents
by: Go, Eun, et al.
Published: (2026)
by: Go, Eun, et al.
Published: (2026)
Conservative Contextual Bandits: Beyond Linear Representations
by: Deb, Rohan, et al.
Published: (2024)
by: Deb, Rohan, et al.
Published: (2024)
Replicable Bandits with UCB based Exploration
by: Deb, Rohan, et al.
Published: (2026)
by: Deb, Rohan, et al.
Published: (2026)
Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms
by: Deb, Rohan, et al.
Published: (2025)
by: Deb, Rohan, et al.
Published: (2025)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
by: Wang, Qi, et al.
Published: (2023)
by: Wang, Qi, et al.
Published: (2023)
A Tractable Inference Perspective of Offline RL
by: Liu, Xuejie, et al.
Published: (2023)
by: Liu, Xuejie, et al.
Published: (2023)
Gradual Fine-Tuning for Flow Matching Models
by: Thorkelsdottir, Gudrun, et al.
Published: (2026)
by: Thorkelsdottir, Gudrun, et al.
Published: (2026)
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
by: Cheng, Jie, et al.
Published: (2024)
by: Cheng, Jie, et al.
Published: (2024)
Loss Gradient Gaussian Width based Generalization and Optimization Guarantees
by: Banerjee, Arindam, et al.
Published: (2024)
by: Banerjee, Arindam, et al.
Published: (2024)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
by: Mark, Max Sobol, et al.
Published: (2024)
by: Mark, Max Sobol, et al.
Published: (2024)
Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model Disentanglement
by: Wang, Zhi, et al.
Published: (2024)
by: Wang, Zhi, et al.
Published: (2024)
Dual Alignment Maximin Optimization for Offline Model-based RL
by: Zhou, Chi, et al.
Published: (2025)
by: Zhou, Chi, et al.
Published: (2025)
Pyramid MoA: A Probabilistic Framework for Cost-Optimized Anytime Inference
by: Khaled, Arindam
Published: (2026)
by: Khaled, Arindam
Published: (2026)
SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation
by: Brita, Catalin E., et al.
Published: (2024)
by: Brita, Catalin E., et al.
Published: (2024)
AdamO: A Collapse-Suppressed Optimizer for Offline RL
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
Optimization for Neural Operators can Benefit from Width
by: Cisneros-Velarde, Pedro, et al.
Published: (2025)
by: Cisneros-Velarde, Pedro, et al.
Published: (2025)
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL
by: Chen, Zhaoyang, et al.
Published: (2025)
by: Chen, Zhaoyang, et al.
Published: (2025)
Diffusion Models as Optimizers for Efficient Planning in Offline RL
by: Huang, Renming, et al.
Published: (2024)
by: Huang, Renming, et al.
Published: (2024)
Offline RL for Adaptive Policy Retrieval in Prior Authorization
by: Sharifullin, Ruslan, et al.
Published: (2026)
by: Sharifullin, Ruslan, et al.
Published: (2026)
Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive Weighting
by: Qiao, Zhongjian, et al.
Published: (2026)
by: Qiao, Zhongjian, et al.
Published: (2026)
Action-Free Offline-to-Online RL via Discretised State Policies
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
Reinformer: Max-Return Sequence Modeling for Offline RL
by: Zhuang, Zifeng, et al.
Published: (2024)
by: Zhuang, Zifeng, et al.
Published: (2024)
Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
by: Landers, Matthew, et al.
Published: (2026)
by: Landers, Matthew, et al.
Published: (2026)
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
by: Soni, Aditya, et al.
Published: (2024)
by: Soni, Aditya, et al.
Published: (2024)
FOSP: Fine-tuning Offline Safe Policy through World Models
by: Cao, Chenyang, et al.
Published: (2024)
by: Cao, Chenyang, et al.
Published: (2024)
Are Expressive Models Truly Necessary for Offline RL?
by: Wang, Guan, et al.
Published: (2024)
by: Wang, Guan, et al.
Published: (2024)
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
by: Mukherjee, Subhojyoti, et al.
Published: (2025)
by: Mukherjee, Subhojyoti, et al.
Published: (2025)
Enhancing Offline Model-Based RL via Active Model Selection: A Bayesian Optimization Perspective
by: Yang, Yu-Wei, et al.
Published: (2025)
by: Yang, Yu-Wei, et al.
Published: (2025)
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
by: Choi, Jinwoo, et al.
Published: (2026)
by: Choi, Jinwoo, et al.
Published: (2026)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
by: He, Longxiang, et al.
Published: (2025)
by: He, Longxiang, et al.
Published: (2025)
Improving Offline RL by Blending Heuristics
by: Geng, Sinong, et al.
Published: (2023)
by: Geng, Sinong, et al.
Published: (2023)
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023)
by: Liu, Yao, et al.
Published: (2023)
How to Make the Gradients Small Privately: Improved Rates for Differentially Private Non-Convex Optimization
by: Lowy, Andrew, et al.
Published: (2024)
by: Lowy, Andrew, et al.
Published: (2024)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
by: Xiao, Wei, et al.
Published: (2025)
by: Xiao, Wei, et al.
Published: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Sketched Adaptive Federated Deep Learning: A Sharp Convergence Analysis
by: Chen, Zhijie, et al.
Published: (2024)
by: Chen, Zhijie, et al.
Published: (2024)
CROP: Conservative Reward for Model-based Offline Policy Optimization
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Scalable Offline Model-Based RL with Action Chunks
by: Park, Kwanyoung, et al.
Published: (2025)
by: Park, Kwanyoung, et al.
Published: (2025)
Similar Items
-
Plan Before You Trade: Inference-Time Optimization for RL Trading Agents
by: Go, Eun, et al.
Published: (2026) -
Conservative Contextual Bandits: Beyond Linear Representations
by: Deb, Rohan, et al.
Published: (2024) -
Replicable Bandits with UCB based Exploration
by: Deb, Rohan, et al.
Published: (2026) -
Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms
by: Deb, Rohan, et al.
Published: (2025) -
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
by: Zhu, Lingwei, et al.
Published: (2025)