DSSP: Diffusion State Space Policy with Full-History Encoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Zhiyuan, Hu, Jianshu, Fang, Han, Jiang, Yunpeng, Huang, Yize, Li, Shujia, Li, Xiao, Ban, Yutong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914584414126080
author Guan, Zhiyuan
Hu, Jianshu
Fang, Han
Jiang, Yunpeng
Huang, Yize
Li, Shujia
Li, Xiao
Ban, Yutong
author_facet Guan, Zhiyuan
Hu, Jianshu
Fang, Han
Jiang, Yunpeng
Huang, Yize
Li, Shujia
Li, Xiao
Ban, Yutong
contents Diffusion-based imitation learning has shown strong promise for robot manipulation. However, most existing policies condition only on the current observation or a short window of recent observations, limiting their ability to resolve history-dependent ambiguities in long-horizon tasks. To address this, we introduce DSSP, a history-conditioned Diffusion State Space Policy that enables efficient, full-history conditioning for robot manipulation. Leveraging the continuous sequence modeling properties of State Space Models (SSMs), our history encoder effectively compresses the entire observation stream into a compact context representation. To ensure this context preserves critical information regarding future state evolution, the encoder is optimized with a dynamics-aware auxiliary training objective. This high-level context representation is then seamlessly fused with recent state observations to form a hierarchical conditioning mechanism for action generation. Furthermore, to maintain architectural consistency and minimize GPU memory overhead, we also instantiate the diffusion backbone itself using an SSM. Extensive experiments across simulation benchmarks and real-world manipulation tasks show that DSSP achieves state-of-the-art performance with a significantly smaller model size, demonstrating superior efficiency of the hierarchical conditioning in capturing crucial information as the history length increases.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14598
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DSSP: Diffusion State Space Policy with Full-History Encoding
Guan, Zhiyuan
Hu, Jianshu
Fang, Han
Jiang, Yunpeng
Huang, Yize
Li, Shujia
Li, Xiao
Ban, Yutong
Robotics
Diffusion-based imitation learning has shown strong promise for robot manipulation. However, most existing policies condition only on the current observation or a short window of recent observations, limiting their ability to resolve history-dependent ambiguities in long-horizon tasks. To address this, we introduce DSSP, a history-conditioned Diffusion State Space Policy that enables efficient, full-history conditioning for robot manipulation. Leveraging the continuous sequence modeling properties of State Space Models (SSMs), our history encoder effectively compresses the entire observation stream into a compact context representation. To ensure this context preserves critical information regarding future state evolution, the encoder is optimized with a dynamics-aware auxiliary training objective. This high-level context representation is then seamlessly fused with recent state observations to form a hierarchical conditioning mechanism for action generation. Furthermore, to maintain architectural consistency and minimize GPU memory overhead, we also instantiate the diffusion backbone itself using an SSM. Extensive experiments across simulation benchmarks and real-world manipulation tasks show that DSSP achieves state-of-the-art performance with a significantly smaller model size, demonstrating superior efficiency of the hierarchical conditioning in capturing crucial information as the history length increases.
title DSSP: Diffusion State Space Policy with Full-History Encoding
topic Robotics
url https://arxiv.org/abs/2605.14598