Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Landers, Matthew, Killian, Taylor W., Hartvigsen, Thomas, Doryab, Afsaneh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAINT: Attention-Based Policies for Discrete Combinatorial Action Spaces
von: Landers, Matthew, et al.
Veröffentlicht: (2025)
von: Landers, Matthew, et al.
Veröffentlicht: (2025)
BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action Spaces
von: Landers, Matthew, et al.
Veröffentlicht: (2024)
von: Landers, Matthew, et al.
Veröffentlicht: (2024)
Coordination Matters: Evaluation of Cooperative Multi-Agent Reinforcement Learning
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2026)
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2026)
Factorized Deep Q-Network for Cooperative Multi-Agent Reinforcement Learning in Victim Tagging
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2025)
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2025)
Action-Free Offline-to-Online RL via Discretised State Policies
von: Neggatu, Natinael Solomon, et al.
Veröffentlicht: (2026)
von: Neggatu, Natinael Solomon, et al.
Veröffentlicht: (2026)
Improving Offline RL by Blending Heuristics
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
Dynamic Neighborhood Construction for Structured Large Discrete Action Spaces
von: Akkerman, Fabian, et al.
Veröffentlicht: (2023)
von: Akkerman, Fabian, et al.
Veröffentlicht: (2023)
Flow Matching for Offline Reinforcement Learning with Discrete Actions
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
Scalable Offline Model-Based RL with Action Chunks
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
Inference Time Policy Optimization for Offline RL with Differentiable World Models
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
Constrained Discrete Diffusion
von: Cardei, Michael, et al.
Veröffentlicht: (2025)
von: Cardei, Michael, et al.
Veröffentlicht: (2025)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
von: Park, Seohong, et al.
Veröffentlicht: (2023)
von: Park, Seohong, et al.
Veröffentlicht: (2023)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
von: Muslimani, Calarina, et al.
Veröffentlicht: (2025)
von: Muslimani, Calarina, et al.
Veröffentlicht: (2025)
An Investigation of Offline Reinforcement Learning in Factorisable Action Spaces
von: Beeson, Alex, et al.
Veröffentlicht: (2024)
von: Beeson, Alex, et al.
Veröffentlicht: (2024)
GEM: Guided Expectation-Maximization for Behavior-Normalized Candidate Action Selection in Offline RL
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL
von: Chen, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoyang, et al.
Veröffentlicht: (2025)
Solving Continual Offline RL through Selective Weights Activation on Aligned Spaces
von: Hu, Jifeng, et al.
Veröffentlicht: (2024)
von: Hu, Jifeng, et al.
Veröffentlicht: (2024)
Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces
von: Gao, Ji, et al.
Veröffentlicht: (2026)
von: Gao, Ji, et al.
Veröffentlicht: (2026)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
von: Zhu, Yuanyang, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanyang, et al.
Veröffentlicht: (2024)
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
Stochastic Q-learning for Large Discrete Action Spaces
von: Fourati, Fares, et al.
Veröffentlicht: (2024)
von: Fourati, Fares, et al.
Veröffentlicht: (2024)
Offline RL for Adaptive Policy Retrieval in Prior Authorization
von: Sharifullin, Ruslan, et al.
Veröffentlicht: (2026)
von: Sharifullin, Ruslan, et al.
Veröffentlicht: (2026)
Dataset Clustering for Improved Offline Policy Learning
von: Wang, Qiang, et al.
Veröffentlicht: (2024)
von: Wang, Qiang, et al.
Veröffentlicht: (2024)
An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model
von: Kang, Enoch H., et al.
Veröffentlicht: (2025)
von: Kang, Enoch H., et al.
Veröffentlicht: (2025)
Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank
von: Zhan, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhan, Wenhao, et al.
Veröffentlicht: (2024)
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
von: Li, Bangzheng, et al.
Veröffentlicht: (2024)
von: Li, Bangzheng, et al.
Veröffentlicht: (2024)
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
von: Choi, Jinwoo, et al.
Veröffentlicht: (2026)
von: Choi, Jinwoo, et al.
Veröffentlicht: (2026)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
von: He, Longxiang, et al.
Veröffentlicht: (2025)
von: He, Longxiang, et al.
Veröffentlicht: (2025)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
von: Killian, Earl
Veröffentlicht: (2026)
von: Killian, Earl
Veröffentlicht: (2026)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
von: K, Swaminathan S, et al.
Veröffentlicht: (2026)
von: K, Swaminathan S, et al.
Veröffentlicht: (2026)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
von: Alvo, Matias, et al.
Veröffentlicht: (2026)
von: Alvo, Matias, et al.
Veröffentlicht: (2026)
From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL
von: Admoni, Sahar, et al.
Veröffentlicht: (2025)
von: Admoni, Sahar, et al.
Veröffentlicht: (2025)
POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition
von: Saito, Yuta, et al.
Veröffentlicht: (2024)
von: Saito, Yuta, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SAINT: Attention-Based Policies for Discrete Combinatorial Action Spaces
von: Landers, Matthew, et al.
Veröffentlicht: (2025) -
BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action Spaces
von: Landers, Matthew, et al.
Veröffentlicht: (2024) -
Coordination Matters: Evaluation of Cooperative Multi-Agent Reinforcement Learning
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2026) -
Factorized Deep Q-Network for Cooperative Multi-Agent Reinforcement Learning in Victim Tagging
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2025) -
Action-Free Offline-to-Online RL via Discretised State Policies
von: Neggatu, Natinael Solomon, et al.
Veröffentlicht: (2026)