VG4D: Vision-Language Model Goes 4D Video Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Deng, Zhichao, Li, Xiangtai, Li, Xia, Tong, Yunhai, Zhao, Shen, Liu, Mengyuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
di: Lu, Haoran, et al.
Pubblicazione: (2026)
di: Lu, Haoran, et al.
Pubblicazione: (2026)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
di: Guo, Jun, et al.
Pubblicazione: (2026)
di: Guo, Jun, et al.
Pubblicazione: (2026)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing
di: Zong, Junyi, et al.
Pubblicazione: (2026)
di: Zong, Junyi, et al.
Pubblicazione: (2026)
Geometry-aware 4D Video Generation for Robot Manipulation
di: Liu, Zeyi, et al.
Pubblicazione: (2025)
di: Liu, Zeyi, et al.
Pubblicazione: (2025)
Unifying 2D and 3D Vision-Language Understanding
di: Jain, Ayush, et al.
Pubblicazione: (2025)
di: Jain, Ayush, et al.
Pubblicazione: (2025)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
di: Babey, Nicholas, et al.
Pubblicazione: (2025)
di: Babey, Nicholas, et al.
Pubblicazione: (2025)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
di: Liu, Yunong, et al.
Pubblicazione: (2024)
di: Liu, Yunong, et al.
Pubblicazione: (2024)
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
di: Xu, Tianling, et al.
Pubblicazione: (2025)
di: Xu, Tianling, et al.
Pubblicazione: (2025)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
di: Li, Ying, et al.
Pubblicazione: (2025)
di: Li, Ying, et al.
Pubblicazione: (2025)
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
di: Sohn, Tin Stribor, et al.
Pubblicazione: (2025)
di: Sohn, Tin Stribor, et al.
Pubblicazione: (2025)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
di: Wen, Yuhang, et al.
Pubblicazione: (2023)
di: Wen, Yuhang, et al.
Pubblicazione: (2023)
Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction
di: Yang, Yiru, et al.
Pubblicazione: (2026)
di: Yang, Yiru, et al.
Pubblicazione: (2026)
RDD4D: 4D Attention-Guided Road Damage Detection And Classification
di: Alkalbani, Asma, et al.
Pubblicazione: (2025)
di: Alkalbani, Asma, et al.
Pubblicazione: (2025)
U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
di: Xu, Xiang, et al.
Pubblicazione: (2025)
di: Xu, Xiang, et al.
Pubblicazione: (2025)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
di: Yan, Xiaoyang, et al.
Pubblicazione: (2026)
di: Yan, Xiaoyang, et al.
Pubblicazione: (2026)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
di: Shen, Boyang, et al.
Pubblicazione: (2026)
di: Shen, Boyang, et al.
Pubblicazione: (2026)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
di: Zhu, He, et al.
Pubblicazione: (2025)
di: Zhu, He, et al.
Pubblicazione: (2025)
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
di: Dong, Xiangyu, et al.
Pubblicazione: (2025)
di: Dong, Xiangyu, et al.
Pubblicazione: (2025)
LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment
di: Peng, Shuaibang, et al.
Pubblicazione: (2026)
di: Peng, Shuaibang, et al.
Pubblicazione: (2026)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
di: Singh, Ishika, et al.
Pubblicazione: (2025)
di: Singh, Ishika, et al.
Pubblicazione: (2025)
3D and 4D World Modeling: A Survey
di: Kong, Lingdong, et al.
Pubblicazione: (2025)
di: Kong, Lingdong, et al.
Pubblicazione: (2025)
4DRaL: Bridging 4D Radar with LiDAR for Place Recognition using Knowledge Distillation
di: Huang, Ningyuan, et al.
Pubblicazione: (2026)
di: Huang, Ningyuan, et al.
Pubblicazione: (2026)
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2024)
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2024)
ModelNet-O: A Large-Scale Synthetic Dataset for Occlusion-Aware Point Cloud Classification
di: Fang, Zhongbin, et al.
Pubblicazione: (2024)
di: Fang, Zhongbin, et al.
Pubblicazione: (2024)
4D-CAAL: 4D Radar-Camera Calibration and Auto-Labeling for Autonomous Driving
di: Yao, Shanliang, et al.
Pubblicazione: (2026)
di: Yao, Shanliang, et al.
Pubblicazione: (2026)
Physically Grounded Vision-Language Models for Robotic Manipulation
di: Gao, Jensen, et al.
Pubblicazione: (2023)
di: Gao, Jensen, et al.
Pubblicazione: (2023)
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
di: Yuan, Chengbo, et al.
Pubblicazione: (2024)
di: Yuan, Chengbo, et al.
Pubblicazione: (2024)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
di: Yang, Sheng, et al.
Pubblicazione: (2025)
di: Yang, Sheng, et al.
Pubblicazione: (2025)
AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
di: Xu, Peng, et al.
Pubblicazione: (2026)
di: Xu, Peng, et al.
Pubblicazione: (2026)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
di: Guo, Jianing, et al.
Pubblicazione: (2025)
di: Guo, Jianing, et al.
Pubblicazione: (2025)
Zero-Shot 3D Visual Grounding from Vision-Language Models
di: Li, Rong, et al.
Pubblicazione: (2025)
di: Li, Rong, et al.
Pubblicazione: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
di: Sarowar, Md Selim, et al.
Pubblicazione: (2026)
di: Sarowar, Md Selim, et al.
Pubblicazione: (2026)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
di: Xu, Mutian, et al.
Pubblicazione: (2026)
di: Xu, Mutian, et al.
Pubblicazione: (2026)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
di: Li, Anqi, et al.
Pubblicazione: (2025)
di: Li, Anqi, et al.
Pubblicazione: (2025)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
di: Zhou, Kaichen, et al.
Pubblicazione: (2026)
di: Zhou, Kaichen, et al.
Pubblicazione: (2026)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
di: Lu, Haoran, et al.
Pubblicazione: (2026) -
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
di: Guo, Jun, et al.
Pubblicazione: (2026) -
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
di: Kim, Jisoo, et al.
Pubblicazione: (2026) -
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
di: Wang, Ziyi, et al.
Pubblicazione: (2025) -
VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing
di: Zong, Junyi, et al.
Pubblicazione: (2026)