ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, You, Chen, Zixuan, Ou, Cunxu, Wang, Wenxuan, Huang, Wenbo, Cao, Lin, Chen, Yangtao, Qiu, Weichao, Quan, Xingyue, Shi, Jieqi, Huo, Jing, Gao, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation
di: Chen, Yangtao, et al.
Pubblicazione: (2024)
di: Chen, Yangtao, et al.
Pubblicazione: (2024)
RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation
di: Chen, Zixuan, et al.
Pubblicazione: (2025)
di: Chen, Zixuan, et al.
Pubblicazione: (2025)
ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon Manipulation
di: Chen, Zixuan, et al.
Pubblicazione: (2025)
di: Chen, Zixuan, et al.
Pubblicazione: (2025)
DeCo: Task Decomposition and Skill Composition for Zero-Shot Generalization in Long-Horizon 3D Manipulation
di: Chen, Zixuan, et al.
Pubblicazione: (2025)
di: Chen, Zixuan, et al.
Pubblicazione: (2025)
RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
di: Chen, Yangtao, et al.
Pubblicazione: (2025)
di: Chen, Yangtao, et al.
Pubblicazione: (2025)
RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation
di: Chen, Zixuan, et al.
Pubblicazione: (2026)
di: Chen, Zixuan, et al.
Pubblicazione: (2026)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
di: Huang, Helong, et al.
Pubblicazione: (2025)
di: Huang, Helong, et al.
Pubblicazione: (2025)
R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
di: Ding, Hongyu, et al.
Pubblicazione: (2025)
di: Ding, Hongyu, et al.
Pubblicazione: (2025)
MoMaStage: Skill-State Graph Guided Planning and Closed-Loop Execution for Long-Horizon Indoor Mobile Manipulation
di: Li, Chenxu, et al.
Pubblicazione: (2026)
di: Li, Chenxu, et al.
Pubblicazione: (2026)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
di: Ma, Chuanhao, et al.
Pubblicazione: (2026)
di: Ma, Chuanhao, et al.
Pubblicazione: (2026)
V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors
di: He, Songjia, et al.
Pubblicazione: (2026)
di: He, Songjia, et al.
Pubblicazione: (2026)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
di: Chen, Jiayi, et al.
Pubblicazione: (2026)
di: Chen, Jiayi, et al.
Pubblicazione: (2026)
INHerit-SG: Incremental Hierarchical Semantic Scene Graphs with RAG-Style Retrieval
di: Fang, YukTungSamuel, et al.
Pubblicazione: (2026)
di: Fang, YukTungSamuel, et al.
Pubblicazione: (2026)
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
di: Cai, Junhao, et al.
Pubblicazione: (2026)
di: Cai, Junhao, et al.
Pubblicazione: (2026)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
di: Zhang, Jiahui, et al.
Pubblicazione: (2025)
di: Zhang, Jiahui, et al.
Pubblicazione: (2025)
ProgVLA: Progress-Aware Robot Manipulation Skill Learning
di: Kim, Seungsu, et al.
Pubblicazione: (2026)
di: Kim, Seungsu, et al.
Pubblicazione: (2026)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
di: Yang, Yandan, et al.
Pubblicazione: (2026)
di: Yang, Yandan, et al.
Pubblicazione: (2026)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
di: Guo, Heyu, et al.
Pubblicazione: (2025)
di: Guo, Heyu, et al.
Pubblicazione: (2025)
NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
di: Chen, Jiahong, et al.
Pubblicazione: (2025)
di: Chen, Jiahong, et al.
Pubblicazione: (2025)
AdaClearGrasp: Learning Adaptive Clearing for Zero-Shot Robust Dexterous Grasping in Densely Cluttered Environments
di: Chen, Zixuan, et al.
Pubblicazione: (2026)
di: Chen, Zixuan, et al.
Pubblicazione: (2026)
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
di: Luo, Bingjun, et al.
Pubblicazione: (2026)
di: Luo, Bingjun, et al.
Pubblicazione: (2026)
VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
di: Gu, Songen, et al.
Pubblicazione: (2026)
di: Gu, Songen, et al.
Pubblicazione: (2026)
VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
di: Bian, Jinyue, et al.
Pubblicazione: (2025)
di: Bian, Jinyue, et al.
Pubblicazione: (2025)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
di: Gao, Juntao, et al.
Pubblicazione: (2025)
di: Gao, Juntao, et al.
Pubblicazione: (2025)
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
di: Sun, Xiaoquan, et al.
Pubblicazione: (2026)
di: Sun, Xiaoquan, et al.
Pubblicazione: (2026)
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
di: Zhou, Hui, et al.
Pubblicazione: (2025)
di: Zhou, Hui, et al.
Pubblicazione: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
di: Ni, Chaojun, et al.
Pubblicazione: (2025)
di: Ni, Chaojun, et al.
Pubblicazione: (2025)
TAIL: A Terrain-Aware Multi-Modal SLAM Dataset for Robot Locomotion in Deformable Granular Environments
di: Yao, Chen, et al.
Pubblicazione: (2024)
di: Yao, Chen, et al.
Pubblicazione: (2024)
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
di: Du, Zhaohui, et al.
Pubblicazione: (2026)
di: Du, Zhaohui, et al.
Pubblicazione: (2026)
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
di: Lian, Shijie, et al.
Pubblicazione: (2026)
di: Lian, Shijie, et al.
Pubblicazione: (2026)
GazeVLA: Learning Human Intention for Robotic Manipulation
di: Li, Chengyang, et al.
Pubblicazione: (2026)
di: Li, Chengyang, et al.
Pubblicazione: (2026)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
di: Jiang, Haoran, et al.
Pubblicazione: (2025)
di: Jiang, Haoran, et al.
Pubblicazione: (2025)
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
di: Li, Wei, et al.
Pubblicazione: (2026)
di: Li, Wei, et al.
Pubblicazione: (2026)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
di: Jiang, Yuming, et al.
Pubblicazione: (2025)
di: Jiang, Yuming, et al.
Pubblicazione: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
di: Yang, Shuai, et al.
Pubblicazione: (2025)
di: Yang, Shuai, et al.
Pubblicazione: (2025)
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
di: Li, Yutai, et al.
Pubblicazione: (2026)
di: Li, Yutai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation
di: Chen, Yangtao, et al.
Pubblicazione: (2024) -
RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation
di: Chen, Zixuan, et al.
Pubblicazione: (2025) -
ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon Manipulation
di: Chen, Zixuan, et al.
Pubblicazione: (2025) -
DeCo: Task Decomposition and Skill Composition for Zero-Shot Generalization in Long-Horizon 3D Manipulation
di: Chen, Zixuan, et al.
Pubblicazione: (2025) -
RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
di: Chen, Yangtao, et al.
Pubblicazione: (2025)