ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Qian, Zhao, Junqiao, Zhou, Hongtu, Yu, Hang, Zhao, Yanping, Ye, Chen, Chen, Guang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914555202895872
author Chen, Qian
Zhao, Junqiao
Zhou, Hongtu
Yu, Hang
Zhao, Yanping
Ye, Chen
Chen, Guang
author_facet Chen, Qian
Zhao, Junqiao
Zhou, Hongtu
Yu, Hang
Zhao, Yanping
Ye, Chen
Chen, Guang
contents Long-horizon, sparse-reward tasks pose a fundamental challenge for reinforcement learning, since single-step TD learning suffers from bootstrapping error accumulation across successive Bellman updates. Actor-critic methods with action chunking address this by operating over temporally extended actions, which reduce the effective horizon, enable fast value backups, and support temporally consistent exploration. However, existing methods rely on a fixed chunk size and therefore cannot adaptively balance reactivity against temporal consistency. A large fixed chunk size reduces responsiveness to new observations, while a small one produces incoherent motions, forcing task-specific tuning of the chunk size. To address this limitation, we propose Adaptive Chunk Size Actor-Critic (ACSAC). ACSAC leverages a causal Transformer critic to evaluate expected returns for action chunks of different sizes. At each chunk boundary, it adaptively selects the chunk size that maximizes the expected return, supporting flexible, state-dependent chunk sizes without task-specific tuning. We prove that the ACSAC Bellman operator is a contraction whose unique fixed point is the action-value function of the adaptive policy. Experiments on OGBench demonstrate that ACSAC achieves state-of-the-art performance on long-horizon, sparse-reward manipulation tasks across both offline RL and offline-to-online RL settings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11009
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
Chen, Qian
Zhao, Junqiao
Zhou, Hongtu
Yu, Hang
Zhao, Yanping
Ye, Chen
Chen, Guang
Machine Learning
Robotics
Long-horizon, sparse-reward tasks pose a fundamental challenge for reinforcement learning, since single-step TD learning suffers from bootstrapping error accumulation across successive Bellman updates. Actor-critic methods with action chunking address this by operating over temporally extended actions, which reduce the effective horizon, enable fast value backups, and support temporally consistent exploration. However, existing methods rely on a fixed chunk size and therefore cannot adaptively balance reactivity against temporal consistency. A large fixed chunk size reduces responsiveness to new observations, while a small one produces incoherent motions, forcing task-specific tuning of the chunk size. To address this limitation, we propose Adaptive Chunk Size Actor-Critic (ACSAC). ACSAC leverages a causal Transformer critic to evaluate expected returns for action chunks of different sizes. At each chunk boundary, it adaptively selects the chunk size that maximizes the expected return, supporting flexible, state-dependent chunk sizes without task-specific tuning. We prove that the ACSAC Bellman operator is a contraction whose unique fixed point is the action-value function of the adaptive policy. Experiments on OGBench demonstrate that ACSAC achieves state-of-the-art performance on long-horizon, sparse-reward manipulation tasks across both offline RL and offline-to-online RL settings.
title ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
topic Machine Learning
Robotics
url https://arxiv.org/abs/2605.11009