Selective Uncertainty Propagation in Offline RL
Fuente:
arXiv
Saved in:
| Main Authors: | Krishnamurthy, Sanath Kumar, Gangwani, Tanmay, Katariya, Sumeet, Kveton, Branislav, Modi, Shrey, Rangi, Anshuka |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
by: Ren, Yinuo, et al.
Published: (2023)
by: Ren, Yinuo, et al.
Published: (2023)
Finite-Time Logarithmic Bayes Regret Upper Bounds
by: Atsidakou, Alexia, et al.
Published: (2023)
by: Atsidakou, Alexia, et al.
Published: (2023)
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
Statistical Guarantees in Synthetic Data through Conformal Adversarial Generation
by: Vishwakarma, Rahul, et al.
Published: (2025)
by: Vishwakarma, Rahul, et al.
Published: (2025)
IITK at SemEval-2024 Task 4: Hierarchical Embeddings for Detection of Persuasion Techniques in Memes
by: Chikoti, Shreenaga, et al.
Published: (2024)
by: Chikoti, Shreenaga, et al.
Published: (2024)
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
by: Mark, Max Sobol, et al.
Published: (2024)
by: Mark, Max Sobol, et al.
Published: (2024)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
by: Fernandez, Nigel, et al.
Published: (2025)
by: Fernandez, Nigel, et al.
Published: (2025)
Is Value Learning Really the Main Bottleneck in Offline RL?
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
Light-Weight Benchmarks Reveal the Hidden Hardware Cost of Zero-Shot Tabular Foundation Models
by: Gangwani, Ishaan, et al.
Published: (2025)
by: Gangwani, Ishaan, et al.
Published: (2025)
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023)
by: Liu, Yao, et al.
Published: (2023)
Agentic Planning with Reasoning for Image Styling via Offline RL
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
Decoupled Prioritized Resampling for Offline RL
by: Yue, Yang, et al.
Published: (2023)
by: Yue, Yang, et al.
Published: (2023)
Augmenting Offline RL with Unlabeled Data
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Heterogeneous Time Constants Improve Stability in Equilibrium Propagation
by: Kubo, Yoshimasa, et al.
Published: (2026)
by: Kubo, Yoshimasa, et al.
Published: (2026)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
by: Kim, Wall, et al.
Published: (2026)
by: Kim, Wall, et al.
Published: (2026)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
by: Kim, Wall, et al.
Published: (2024)
by: Kim, Wall, et al.
Published: (2024)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
by: Su, Jianhai, et al.
Published: (2025)
by: Su, Jianhai, et al.
Published: (2025)
A Tractable Inference Perspective of Offline RL
by: Liu, Xuejie, et al.
Published: (2023)
by: Liu, Xuejie, et al.
Published: (2023)
Design Considerations in Offline Preference-based RL
by: Agarwal, Alekh, et al.
Published: (2025)
by: Agarwal, Alekh, et al.
Published: (2025)
Are Expressive Models Truly Necessary for Offline RL?
by: Wang, Guan, et al.
Published: (2024)
by: Wang, Guan, et al.
Published: (2024)
OGBench: Benchmarking Offline Goal-Conditioned RL
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
GLIDE-RL: Grounded Language Instruction through DEmonstration in RL
by: Kharyal, Chaitanya, et al.
Published: (2024)
by: Kharyal, Chaitanya, et al.
Published: (2024)
Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams
by: Bansal, Aayam, et al.
Published: (2025)
by: Bansal, Aayam, et al.
Published: (2025)
Language-Model Prior Overcomes Cold-Start Items
by: Wang, Shiyu, et al.
Published: (2024)
by: Wang, Shiyu, et al.
Published: (2024)
Offline Multi-task Transfer RL with Representational Penalization
by: Bose, Avinandan, et al.
Published: (2024)
by: Bose, Avinandan, et al.
Published: (2024)
Yes, Q-learning Helps Offline In-Context RL
by: Tarasov, Denis, et al.
Published: (2025)
by: Tarasov, Denis, et al.
Published: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
Scalable Offline Model-Based RL with Action Chunks
by: Park, Kwanyoung, et al.
Published: (2025)
by: Park, Kwanyoung, et al.
Published: (2025)
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
by: Beck, Jacob
Published: (2025)
by: Beck, Jacob
Published: (2025)
Dual Alignment Maximin Optimization for Offline Model-based RL
by: Zhou, Chi, et al.
Published: (2025)
by: Zhou, Chi, et al.
Published: (2025)
Integrating Domain Knowledge for handling Limited Data in Offline RL
by: Gangopadhyay, Briti, et al.
Published: (2024)
by: Gangopadhyay, Briti, et al.
Published: (2024)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
Language-Conditioned Offline RL for Multi-Robot Navigation
by: Morad, Steven, et al.
Published: (2024)
by: Morad, Steven, et al.
Published: (2024)
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation
by: Batra, Sumeet, et al.
Published: (2024)
by: Batra, Sumeet, et al.
Published: (2024)
Action-Free Offline-to-Online RL via Discretised State Policies
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
Similar Items
-
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026) -
Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
by: Ren, Yinuo, et al.
Published: (2023) -
Finite-Time Logarithmic Bayes Regret Upper Bounds
by: Atsidakou, Alexia, et al.
Published: (2023) -
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
by: Kveton, Branislav, et al.
Published: (2026) -
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)