Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Jinzong, Huang, Wei, Zhang, Jianshu, Chen, Zhuo, Yuan, Xinzhe, Gu, Qinying, Jiang, Zhaohui, Ye, Nanyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
by: Yuan, Xinzhe, et al.
Published: (2026)
by: Yuan, Xinzhe, et al.
Published: (2026)
LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective Experimentation
by: Chen, Zhuo, et al.
Published: (2026)
by: Chen, Zhuo, et al.
Published: (2026)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Expectation Consistency Loss: Rethink Confidence Calibration under Covariate Shift
by: Dong, Jinzong, et al.
Published: (2026)
by: Dong, Jinzong, et al.
Published: (2026)
Combining Priors with Experience: Confidence Calibration Based on Binomial Process Modeling
by: Dong, Jinzong, et al.
Published: (2024)
by: Dong, Jinzong, et al.
Published: (2024)
Rethinking ASTE: A Minimalist Tagging Scheme Alongside Contrastive Learning
by: Sun, Qiao, et al.
Published: (2024)
by: Sun, Qiao, et al.
Published: (2024)
Flow Actor-Critic for Offline Reinforcement Learning
by: Chae, Jongseong, et al.
Published: (2026)
by: Chae, Jongseong, et al.
Published: (2026)
Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration
by: Xian, Mingtao, et al.
Published: (2026)
by: Xian, Mingtao, et al.
Published: (2026)
Enhancing Nursing and Elderly Care with Large Language Models: An AI-Driven Framework
by: Sun, Qiao, et al.
Published: (2024)
by: Sun, Qiao, et al.
Published: (2024)
MiniConGTS: A Near Ultimate Minimalist Contrastive Grid Tagging Scheme for Aspect Sentiment Triplet Extraction
by: Sun, Qiao, et al.
Published: (2024)
by: Sun, Qiao, et al.
Published: (2024)
Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models
by: Zhu, Lin, et al.
Published: (2025)
by: Zhu, Lin, et al.
Published: (2025)
$Δ\mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
by: Zhu, Lin, et al.
Published: (2025)
by: Zhu, Lin, et al.
Published: (2025)
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning
by: Wei, Honghao, et al.
Published: (2024)
by: Wei, Honghao, et al.
Published: (2024)
SelfBC: Self Behavior Cloning for Offline Reinforcement Learning
by: Liu, Shirong, et al.
Published: (2024)
by: Liu, Shirong, et al.
Published: (2024)
TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection
by: Yang, Yifeng, et al.
Published: (2026)
by: Yang, Yifeng, et al.
Published: (2026)
Offline Actor-Critic Reinforcement Learning Scales to Large Models
by: Springenberg, Jost Tobias, et al.
Published: (2024)
by: Springenberg, Jost Tobias, et al.
Published: (2024)
OODD: Test-time Out-of-Distribution Detection with Dynamic Dictionary
by: Yang, Yifeng, et al.
Published: (2025)
by: Yang, Yifeng, et al.
Published: (2025)
Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces
by: Gao, Ji, et al.
Published: (2026)
by: Gao, Ji, et al.
Published: (2026)
Learning Partial Action Replacement in Offline MARL
by: Jin, Yue, et al.
Published: (2026)
by: Jin, Yue, et al.
Published: (2026)
Actor-Critic Reinforcement Learning with Phased Actor
by: Wu, Ruofan, et al.
Published: (2024)
by: Wu, Ruofan, et al.
Published: (2024)
Effective Reinforcement Learning Control using Conservative Soft Actor-Critic
by: Shang, Zhiwei, et al.
Published: (2025)
by: Shang, Zhiwei, et al.
Published: (2025)
Risk-sensitive Actor-Critic with Static Spectral Risk Measures for Online and Offline Reinforcement Learning
by: Moghimi, Mehrdad, et al.
Published: (2025)
by: Moghimi, Mehrdad, et al.
Published: (2025)
Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward
by: Yang, Jiarui, et al.
Published: (2025)
by: Yang, Jiarui, et al.
Published: (2025)
Actor-Critic Pretraining for Proximal Policy Optimization
by: Kernbach, Andreas, et al.
Published: (2026)
by: Kernbach, Andreas, et al.
Published: (2026)
Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning
by: Fang, Linjiajie, et al.
Published: (2024)
by: Fang, Linjiajie, et al.
Published: (2024)
A Pontryagin Method of Model-based Reinforcement Learning via Hamiltonian Actor-Critic
by: Gu, Chengyang, et al.
Published: (2026)
by: Gu, Chengyang, et al.
Published: (2026)
Understanding Behavior Cloning with Action Quantization
by: Cao, Haoqun, et al.
Published: (2026)
by: Cao, Haoqun, et al.
Published: (2026)
Counterfactual Behavior Cloning: Offline Imitation Learning from Imperfect Human Demonstrations
by: Sagheb, Shahabedin, et al.
Published: (2025)
by: Sagheb, Shahabedin, et al.
Published: (2025)
State-Novelty Guided Action Persistence in Deep Reinforcement Learning
by: Hu, Jianshu, et al.
Published: (2024)
by: Hu, Jianshu, et al.
Published: (2024)
SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
Quantum Advantage Actor-Critic for Reinforcement Learning
by: Kölle, Michael, et al.
Published: (2024)
by: Kölle, Michael, et al.
Published: (2024)
FORLER: Federated Offline Reinforcement Learning with Q-Ensemble and Actor Rectification
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
Flow Matching for Offline Reinforcement Learning with Discrete Actions
by: Khan, Fairoz Nower, et al.
Published: (2026)
by: Khan, Fairoz Nower, et al.
Published: (2026)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning
by: Chen, Haohui, et al.
Published: (2024)
by: Chen, Haohui, et al.
Published: (2024)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
by: Sun, Qiao, et al.
Published: (2025)
by: Sun, Qiao, et al.
Published: (2025)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
by: Yang, Tong, et al.
Published: (2023)
by: Yang, Tong, et al.
Published: (2023)
Pretraining in Actor-Critic Reinforcement Learning for Robot Locomotion
by: Fan, Jiale, et al.
Published: (2025)
by: Fan, Jiale, et al.
Published: (2025)
Partial Action Replacement: Tackling Distribution Shift in Offline MARL
by: Jin, Yue, et al.
Published: (2025)
by: Jin, Yue, et al.
Published: (2025)
Similar Items
-
Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
by: Yuan, Xinzhe, et al.
Published: (2026) -
LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective Experimentation
by: Chen, Zhuo, et al.
Published: (2026) -
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
by: Huang, Wei, et al.
Published: (2025) -
Expectation Consistency Loss: Rethink Confidence Calibration under Covariate Shift
by: Dong, Jinzong, et al.
Published: (2026) -
Combining Priors with Experience: Confidence Calibration Based on Binomial Process Modeling
by: Dong, Jinzong, et al.
Published: (2024)