ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Hang, Zhang, Di, Du, Qiwei, Zhao, Yanping, Zhang, Hai, Chen, Guang, Veas, Eduardo E., Zhao, Junqiao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915677770612736
author Yu, Hang
Zhang, Di
Du, Qiwei
Zhao, Yanping
Zhang, Hai
Chen, Guang
Veas, Eduardo E.
Zhao, Junqiao
author_facet Yu, Hang
Zhang, Di
Du, Qiwei
Zhao, Yanping
Zhang, Hai
Chen, Guang
Veas, Eduardo E.
Zhao, Junqiao
contents Offline reinforcement learning (RL) enables agents to learn optimal policies from pre-collected datasets. However, datasets containing suboptimal and fragmented trajectories present challenges for reward propagation, resulting in inaccurate value estimation and degraded policy performance. While trajectory stitching via generative models offers a promising solution, existing augmentation methods frequently produce trajectories that are either confined to the support of the behavior policy or violate the underlying dynamics, thereby limiting their effectiveness for policy improvement. We propose ASTRO, a data augmentation framework that generates distributionally novel and dynamics-consistent trajectories for offline RL. ASTRO first learns a temporal-distance representation to identify distinct and reachable stitch targets. We then employ a dynamics-guided stitch planner that adaptively generates connecting action sequences via Rollout Deviation Feedback, defined as the gap between target state sequence and the actual arrived state sequence by executing predicted actions, to improve trajectory stitching's feasibility and reachability. This approach facilitates effective augmentation through stitching and ultimately enhances policy learning. ASTRO outperforms prior offline RL augmentation methods across various algorithms, achieving notable performance gain on the challenging OGBench suite and demonstrating consistent improvements on standard offline RL benchmarks such as D4RL.
format Preprint
id arxiv_https___arxiv_org_abs_2511_23442
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts
Yu, Hang
Zhang, Di
Du, Qiwei
Zhao, Yanping
Zhang, Hai
Chen, Guang
Veas, Eduardo E.
Zhao, Junqiao
Machine Learning
Artificial Intelligence
Offline reinforcement learning (RL) enables agents to learn optimal policies from pre-collected datasets. However, datasets containing suboptimal and fragmented trajectories present challenges for reward propagation, resulting in inaccurate value estimation and degraded policy performance. While trajectory stitching via generative models offers a promising solution, existing augmentation methods frequently produce trajectories that are either confined to the support of the behavior policy or violate the underlying dynamics, thereby limiting their effectiveness for policy improvement. We propose ASTRO, a data augmentation framework that generates distributionally novel and dynamics-consistent trajectories for offline RL. ASTRO first learns a temporal-distance representation to identify distinct and reachable stitch targets. We then employ a dynamics-guided stitch planner that adaptively generates connecting action sequences via Rollout Deviation Feedback, defined as the gap between target state sequence and the actual arrived state sequence by executing predicted actions, to improve trajectory stitching's feasibility and reachability. This approach facilitates effective augmentation through stitching and ultimately enhances policy learning. ASTRO outperforms prior offline RL augmentation methods across various algorithms, achieving notable performance gain on the challenging OGBench suite and demonstrating consistent improvements on standard offline RL benchmarks such as D4RL.
title ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.23442