Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Babey, Nicholas, Gu, Tiffany, Li, Yiheng, Meo, Cristian, Zhu, Kevin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
von: Wen, Bowen, et al.
Veröffentlicht: (2023)
von: Wen, Bowen, et al.
Veröffentlicht: (2023)
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
von: Wang, Guodong, et al.
Veröffentlicht: (2026)
von: Wang, Guodong, et al.
Veröffentlicht: (2026)
Acoustic-based 3D Human Pose Estimation Robust to Human Position
von: Oumi, Yusuke, et al.
Veröffentlicht: (2024)
von: Oumi, Yusuke, et al.
Veröffentlicht: (2024)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
Universal Actions for Enhanced Embodied Foundation Models
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
VG4D: Vision-Language Model Goes 4D Video Recognition
von: Deng, Zhichao, et al.
Veröffentlicht: (2024)
von: Deng, Zhichao, et al.
Veröffentlicht: (2024)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
von: Xu, Peng, et al.
Veröffentlicht: (2026)
von: Xu, Peng, et al.
Veröffentlicht: (2026)
3D-VLA: A 3D Vision-Language-Action Generative World Model
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
Detection, Recognition and Pose Estimation of Tabletop Objects
von: Nirgude, Sanjuksha, et al.
Veröffentlicht: (2024)
von: Nirgude, Sanjuksha, et al.
Veröffentlicht: (2024)
Physically Grounded Vision-Language Models for Robotic Manipulation
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
Playing to Vision Foundation Model's Strengths in Stereo Matching
von: Liu, Chuang-Wei, et al.
Veröffentlicht: (2024)
von: Liu, Chuang-Wei, et al.
Veröffentlicht: (2024)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
von: Xu, Tianling, et al.
Veröffentlicht: (2025)
von: Xu, Tianling, et al.
Veröffentlicht: (2025)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
von: Kim, Jisoo, et al.
Veröffentlicht: (2026)
von: Kim, Jisoo, et al.
Veröffentlicht: (2026)
Multi-view Pose Fusion for Occlusion-Aware 3D Human Pose Estimation
von: Bragagnolo, Laura, et al.
Veröffentlicht: (2024)
von: Bragagnolo, Laura, et al.
Veröffentlicht: (2024)
SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation
von: Chen, Ziyi, et al.
Veröffentlicht: (2025)
von: Chen, Ziyi, et al.
Veröffentlicht: (2025)
FlyPose: Towards Robust Human Pose Estimation From Aerial Views
von: Farooq, Hassaan, et al.
Veröffentlicht: (2026)
von: Farooq, Hassaan, et al.
Veröffentlicht: (2026)
Any6D: Model-free 6D Pose Estimation of Novel Objects
von: Lee, Taeyeop, et al.
Veröffentlicht: (2025)
von: Lee, Taeyeop, et al.
Veröffentlicht: (2025)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
von: Yang, Zhenjie, et al.
Veröffentlicht: (2025)
von: Yang, Zhenjie, et al.
Veröffentlicht: (2025)
SonarSweep: Fusing Sonar and Vision for Robust 3D Reconstruction via Plane Sweeping
von: Chen, Lingpeng, et al.
Veröffentlicht: (2025)
von: Chen, Lingpeng, et al.
Veröffentlicht: (2025)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
von: Li, Anqi, et al.
Veröffentlicht: (2025)
von: Li, Anqi, et al.
Veröffentlicht: (2025)
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
von: Yajima, Masaru, et al.
Veröffentlicht: (2025)
von: Yajima, Masaru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025) -
FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
von: Wen, Bowen, et al.
Veröffentlicht: (2023) -
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
von: Wang, Guodong, et al.
Veröffentlicht: (2026) -
Acoustic-based 3D Human Pose Estimation Robust to Human Position
von: Oumi, Yusuke, et al.
Veröffentlicht: (2024) -
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)