ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Chuanhao, Zhou, Hanyu, Peng, Shihan, Li, Yan, Gu, Tao, Yan, Luxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
von: Zhou, Hanyu, et al.
Veröffentlicht: (2026)
von: Zhou, Hanyu, et al.
Veröffentlicht: (2026)
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
von: Yue, Lu, et al.
Veröffentlicht: (2025)
von: Yue, Lu, et al.
Veröffentlicht: (2025)
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection
von: Zhou, Hanyu, et al.
Veröffentlicht: (2024)
von: Zhou, Hanyu, et al.
Veröffentlicht: (2024)
ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
von: Wang, Haonan, et al.
Veröffentlicht: (2026)
von: Wang, Haonan, et al.
Veröffentlicht: (2026)
3DGS-Calib: 3D Gaussian Splatting for Multimodal SpatioTemporal Calibration
von: Herau, Quentin, et al.
Veröffentlicht: (2024)
von: Herau, Quentin, et al.
Veröffentlicht: (2024)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)
Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation
von: Ye, Guo, et al.
Veröffentlicht: (2025)
von: Ye, Guo, et al.
Veröffentlicht: (2025)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
SpatioTemporal Difference Network for Video Depth Super-Resolution
von: Wang, Zhengxue, et al.
Veröffentlicht: (2025)
von: Wang, Zhengxue, et al.
Veröffentlicht: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
Think Proprioceptively: Embodied Visual Reasoning for VLA Manipulation
von: Wang, Fangyuan, et al.
Veröffentlicht: (2026)
von: Wang, Fangyuan, et al.
Veröffentlicht: (2026)
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
von: Huang, Wenlong, et al.
Veröffentlicht: (2024)
von: Huang, Wenlong, et al.
Veröffentlicht: (2024)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
von: Yu, Jiawen, et al.
Veröffentlicht: (2025)
von: Yu, Jiawen, et al.
Veröffentlicht: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
Seeing Motion at Nighttime with an Event Camera
von: Liu, Haoyue, et al.
Veröffentlicht: (2024)
von: Liu, Haoyue, et al.
Veröffentlicht: (2024)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
von: Chopra, Samarth, et al.
Veröffentlicht: (2025)
von: Chopra, Samarth, et al.
Veröffentlicht: (2025)
ST-GS: Vision-Based 3D Semantic Occupancy Prediction with Spatial-Temporal Gaussian Splatting
von: Yan, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Yan, Xiaoyang, et al.
Veröffentlicht: (2025)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
RoboPearls: Editable Video Simulation for Robot Manipulation
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
von: Wang, Haonan, et al.
Veröffentlicht: (2026)
von: Wang, Haonan, et al.
Veröffentlicht: (2026)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
von: Cui, Can, et al.
Veröffentlicht: (2025)
von: Cui, Can, et al.
Veröffentlicht: (2025)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025) -
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
von: Zhou, Hanyu, et al.
Veröffentlicht: (2026) -
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
von: Wang, Haonan, et al.
Veröffentlicht: (2025) -
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
von: Yue, Lu, et al.
Veröffentlicht: (2025) -
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
von: Zhou, Hanyu, et al.
Veröffentlicht: (2025)