Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908827053457408 |
|---|---|
| author | Zhan, Simon Sinong Wu, Qingyuan Wang, Philip Yang, Frank Shi, Xiangyu Huang, Chao Zhu, Qi |
| author_facet | Zhan, Simon Sinong Wu, Qingyuan Wang, Philip Yang, Frank Shi, Xiangyu Huang, Chao Zhu, Qi |
| contents | Offline-to-online deployment of reinforcement-learning (RL) agents must bridge two gaps: (1) the sim-to-real gap, where real systems add latency and other imperfections not present in simulation, and (2) the interaction gap, where policies trained purely offline face out-of-distribution states during online execution because gathering new interaction data is costly or risky. Agents therefore have to generalize from static, delay-free datasets to dynamic, delay-prone environments. Standard offline RL learns from delay-free logs yet must act under delays that break the Markov assumption and hurt performance. We introduce DT-CORL (Delay-Transformer belief policy Constrained Offline RL), an offline-RL framework built to cope with delayed dynamics at deployment. DT-CORL (i) produces delay-robust actions with a transformer-based belief predictor even though it never sees delayed observations during training, and (ii) is markedly more sample-efficient than naïve history-augmentation baselines. Experiments on D4RL benchmarks with several delay settings show that DT-CORL consistently outperforms both history-augmentation and vanilla belief-based methods, narrowing the sim-to-real latency gap while preserving data efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_00131 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization Zhan, Simon Sinong Wu, Qingyuan Wang, Philip Yang, Frank Shi, Xiangyu Huang, Chao Zhu, Qi Machine Learning Artificial Intelligence Offline-to-online deployment of reinforcement-learning (RL) agents must bridge two gaps: (1) the sim-to-real gap, where real systems add latency and other imperfections not present in simulation, and (2) the interaction gap, where policies trained purely offline face out-of-distribution states during online execution because gathering new interaction data is costly or risky. Agents therefore have to generalize from static, delay-free datasets to dynamic, delay-prone environments. Standard offline RL learns from delay-free logs yet must act under delays that break the Markov assumption and hurt performance. We introduce DT-CORL (Delay-Transformer belief policy Constrained Offline RL), an offline-RL framework built to cope with delayed dynamics at deployment. DT-CORL (i) produces delay-robust actions with a transformer-based belief predictor even though it never sees delayed observations during training, and (ii) is markedly more sample-efficient than naïve history-augmentation baselines. Experiments on D4RL benchmarks with several delay settings show that DT-CORL consistently outperforms both history-augmentation and vanilla belief-based methods, narrowing the sim-to-real latency gap while preserving data efficiency. |
| title | Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2506.00131 |