ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Chuanhao, Zhou, Hanyu, Peng, Shihan, Li, Yan, Gu, Tao, Yan, Luxin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
by: Zhou, Hanyu, et al.
Published: (2026)
by: Zhou, Hanyu, et al.
Published: (2026)
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
by: Yue, Lu, et al.
Published: (2025)
by: Yue, Lu, et al.
Published: (2025)
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection
by: Zhou, Hanyu, et al.
Published: (2024)
by: Zhou, Hanyu, et al.
Published: (2024)
ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
by: Wang, Haonan, et al.
Published: (2026)
by: Wang, Haonan, et al.
Published: (2026)
3DGS-Calib: 3D Gaussian Splatting for Multimodal SpatioTemporal Calibration
by: Herau, Quentin, et al.
Published: (2024)
by: Herau, Quentin, et al.
Published: (2024)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
by: Wang, Yiru, et al.
Published: (2026)
by: Wang, Yiru, et al.
Published: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025)
by: Hanyu, Taisei, et al.
Published: (2025)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
by: Guo, Heyu, et al.
Published: (2025)
by: Guo, Heyu, et al.
Published: (2025)
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation
by: Ye, Guo, et al.
Published: (2025)
by: Ye, Guo, et al.
Published: (2025)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
by: Li, Runhao, et al.
Published: (2025)
by: Li, Runhao, et al.
Published: (2025)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
by: Wu, Zhenyu, et al.
Published: (2025)
by: Wu, Zhenyu, et al.
Published: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
by: Jiang, Yuming, et al.
Published: (2025)
by: Jiang, Yuming, et al.
Published: (2025)
Think Proprioceptively: Embodied Visual Reasoning for VLA Manipulation
by: Wang, Fangyuan, et al.
Published: (2026)
by: Wang, Fangyuan, et al.
Published: (2026)
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
by: Huang, Wenlong, et al.
Published: (2024)
by: Huang, Wenlong, et al.
Published: (2024)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
by: Han, ByungOk, et al.
Published: (2024)
by: Han, ByungOk, et al.
Published: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
by: Yu, Jiawen, et al.
Published: (2025)
by: Yu, Jiawen, et al.
Published: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
Seeing Motion at Nighttime with an Event Camera
by: Liu, Haoyue, et al.
Published: (2024)
by: Liu, Haoyue, et al.
Published: (2024)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
by: Chopra, Samarth, et al.
Published: (2025)
by: Chopra, Samarth, et al.
Published: (2025)
ST-GS: Vision-Based 3D Semantic Occupancy Prediction with Spatial-Temporal Gaussian Splatting
by: Yan, Xiaoyang, et al.
Published: (2025)
by: Yan, Xiaoyang, et al.
Published: (2025)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
by: Yang, Yandan, et al.
Published: (2026)
by: Yang, Yandan, et al.
Published: (2026)
RoboPearls: Editable Video Simulation for Robot Manipulation
by: Tang, Tao, et al.
Published: (2025)
by: Tang, Tao, et al.
Published: (2025)
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
by: Wang, Haonan, et al.
Published: (2026)
by: Wang, Haonan, et al.
Published: (2026)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
by: Jiang, Haoran, et al.
Published: (2025)
by: Jiang, Haoran, et al.
Published: (2025)
Similar Items
-
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025) -
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
by: Zhou, Hanyu, et al.
Published: (2026) -
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025) -
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
by: Yue, Lu, et al.
Published: (2025) -
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025)