JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Miao, Shangchen, Feng, Ningya, Wu, Jialong, Lin, Ye, He, Xu, Li, Dong, Long, Mingsheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
iVideoGPT: Interactive VideoGPTs are Scalable World Models
di: Wu, Jialong, et al.
Pubblicazione: (2024)
di: Wu, Jialong, et al.
Pubblicazione: (2024)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
di: Sun, Jingwen, et al.
Pubblicazione: (2026)
di: Sun, Jingwen, et al.
Pubblicazione: (2026)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
di: Guo, Wenxuan, et al.
Pubblicazione: (2026)
di: Guo, Wenxuan, et al.
Pubblicazione: (2026)
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
di: Ni, Chaojun, et al.
Pubblicazione: (2025)
di: Ni, Chaojun, et al.
Pubblicazione: (2025)
Vid2World: Crafting Video Diffusion Models to Interactive World Models
di: Huang, Siqiao, et al.
Pubblicazione: (2025)
di: Huang, Siqiao, et al.
Pubblicazione: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
di: Liu, Jiahang, et al.
Pubblicazione: (2025)
di: Liu, Jiahang, et al.
Pubblicazione: (2025)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
di: Fan, Xianzhe, et al.
Pubblicazione: (2026)
di: Fan, Xianzhe, et al.
Pubblicazione: (2026)
A Pragmatic VLA Foundation Model
di: Wu, Wei, et al.
Pubblicazione: (2026)
di: Wu, Wei, et al.
Pubblicazione: (2026)
TrackVLA: Embodied Visual Tracking in the Wild
di: Wang, Shaoan, et al.
Pubblicazione: (2025)
di: Wang, Shaoan, et al.
Pubblicazione: (2025)
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
di: jia, Feiyang, et al.
Pubblicazione: (2026)
di: jia, Feiyang, et al.
Pubblicazione: (2026)
GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
di: Qian, Jingjing, et al.
Pubblicazione: (2025)
di: Qian, Jingjing, et al.
Pubblicazione: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
di: Li, Qiwei, et al.
Pubblicazione: (2026)
di: Li, Qiwei, et al.
Pubblicazione: (2026)
Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation
di: Ye, Guo, et al.
Pubblicazione: (2025)
di: Ye, Guo, et al.
Pubblicazione: (2025)
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
di: Yu, Jiawen, et al.
Pubblicazione: (2025)
di: Yu, Jiawen, et al.
Pubblicazione: (2025)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
di: Li, Yixuan, et al.
Pubblicazione: (2025)
di: Li, Yixuan, et al.
Pubblicazione: (2025)
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving
di: Huang, Yuzhou, et al.
Pubblicazione: (2026)
di: Huang, Yuzhou, et al.
Pubblicazione: (2026)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
di: Liang, Wenqi, et al.
Pubblicazione: (2025)
di: Liang, Wenqi, et al.
Pubblicazione: (2025)
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving
di: Liu, Qiqi, et al.
Pubblicazione: (2026)
di: Liu, Qiqi, et al.
Pubblicazione: (2026)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
di: Ye, Angen, et al.
Pubblicazione: (2025)
di: Ye, Angen, et al.
Pubblicazione: (2025)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
di: Zhang, Wenyao, et al.
Pubblicazione: (2025)
di: Zhang, Wenyao, et al.
Pubblicazione: (2025)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
di: Guo, Heyu, et al.
Pubblicazione: (2025)
di: Guo, Heyu, et al.
Pubblicazione: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
di: Wu, Zhenyu, et al.
Pubblicazione: (2025)
di: Wu, Zhenyu, et al.
Pubblicazione: (2025)
GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies
di: Neau, Maëlic, et al.
Pubblicazione: (2025)
di: Neau, Maëlic, et al.
Pubblicazione: (2025)
LLaDA-VLA: Vision Language Diffusion Action Models
di: Wen, Yuqing, et al.
Pubblicazione: (2025)
di: Wen, Yuqing, et al.
Pubblicazione: (2025)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
di: Wen, Junjie, et al.
Pubblicazione: (2025)
di: Wen, Junjie, et al.
Pubblicazione: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
di: Shen, Yichao, et al.
Pubblicazione: (2025)
di: Shen, Yichao, et al.
Pubblicazione: (2025)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
di: Xu, Tianyu, et al.
Pubblicazione: (2025)
di: Xu, Tianyu, et al.
Pubblicazione: (2025)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
di: Li, Yongkang, et al.
Pubblicazione: (2026)
di: Li, Yongkang, et al.
Pubblicazione: (2026)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
di: Jiang, Haoran, et al.
Pubblicazione: (2025)
di: Jiang, Haoran, et al.
Pubblicazione: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
di: Li, Wei, et al.
Pubblicazione: (2025)
di: Li, Wei, et al.
Pubblicazione: (2025)
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
di: Yu, Bin, et al.
Pubblicazione: (2026)
di: Yu, Bin, et al.
Pubblicazione: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models
di: Chen, Canyu, et al.
Pubblicazione: (2026)
di: Chen, Canyu, et al.
Pubblicazione: (2026)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
di: Li, Runhao, et al.
Pubblicazione: (2025)
di: Li, Runhao, et al.
Pubblicazione: (2025)
VLANeXt: Recipes for Building Strong VLA Models
di: Wu, Xiao-Ming, et al.
Pubblicazione: (2026)
di: Wu, Xiao-Ming, et al.
Pubblicazione: (2026)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
di: Zhang, Qiyao, et al.
Pubblicazione: (2026)
di: Zhang, Qiyao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
iVideoGPT: Interactive VideoGPTs are Scalable World Models
di: Wu, Jialong, et al.
Pubblicazione: (2024) -
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
di: Sun, Jingwen, et al.
Pubblicazione: (2026) -
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
di: Guo, Wenxuan, et al.
Pubblicazione: (2026) -
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
di: Ni, Chaojun, et al.
Pubblicazione: (2025) -
Vid2World: Crafting Video Diffusion Models to Interactive World Models
di: Huang, Siqiao, et al.
Pubblicazione: (2025)