CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Minelli, Giovanni, Turrisi, Giulio, Barasuol, Victor, Semini, Claudio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908872770322432
author Minelli, Giovanni
Turrisi, Giulio
Barasuol, Victor
Semini, Claudio
author_facet Minelli, Giovanni
Turrisi, Giulio
Barasuol, Victor
Semini, Claudio
contents Learning robotic manipulation policies through supervised learning from demonstrations remains challenging when policies encounter execution variations not explicitly covered during training. While incorporating historical context through attention mechanisms can improve robustness, standard approaches process all past states in a sequence without explicitly modeling the temporal structure that demonstrations may include, such as failure and recovery patterns. We propose a Cross-State Transition Attention Transformer that employs a novel State Transition Attention (STA) mechanism to modulate standard attention weights based on learned state evolution patterns, enabling policies to better adapt their behavior based on execution history. Our approach combines this structured attention with temporal masking during training, where visual information is randomly removed from recent timesteps to encourage temporal reasoning from historical context. Evaluation in simulation shows that STA consistently outperforms standard attention approach and temporal modeling methods like TCN and LSTM networks, achieving more than 2x improvement over cross-attention on precision-critical tasks. The source code and data can be accessed at https://github.com/iit-DLSLab/croSTAta
format Preprint
id arxiv_https___arxiv_org_abs_2510_00726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation
Minelli, Giovanni
Turrisi, Giulio
Barasuol, Victor
Semini, Claudio
Robotics
Artificial Intelligence
Machine Learning
Learning robotic manipulation policies through supervised learning from demonstrations remains challenging when policies encounter execution variations not explicitly covered during training. While incorporating historical context through attention mechanisms can improve robustness, standard approaches process all past states in a sequence without explicitly modeling the temporal structure that demonstrations may include, such as failure and recovery patterns. We propose a Cross-State Transition Attention Transformer that employs a novel State Transition Attention (STA) mechanism to modulate standard attention weights based on learned state evolution patterns, enabling policies to better adapt their behavior based on execution history. Our approach combines this structured attention with temporal masking during training, where visual information is randomly removed from recent timesteps to encourage temporal reasoning from historical context. Evaluation in simulation shows that STA consistently outperforms standard attention approach and temporal modeling methods like TCN and LSTM networks, achieving more than 2x improvement over cross-attention on precision-critical tasks. The source code and data can be accessed at https://github.com/iit-DLSLab/croSTAta
title CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.00726