High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Yongpeng, Yamakawa, Yuji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion
von: Gregorek, Jakub, et al.
Veröffentlicht: (2026)
von: Gregorek, Jakub, et al.
Veröffentlicht: (2026)
ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control
von: Wang, Haosen, et al.
Veröffentlicht: (2026)
von: Wang, Haosen, et al.
Veröffentlicht: (2026)
MLFM: Multi-Layered Feature Maps for Richer Language Understanding in Zero-Shot Semantic Navigation
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2025)
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2025)
DriveVA: Video Action Models are Zero-Shot Drivers
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
Improving Zero-Shot ObjectNav with Generative Communication
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2024)
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2024)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
von: Chen, Kehan, et al.
Veröffentlicht: (2024)
von: Chen, Kehan, et al.
Veröffentlicht: (2024)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
Zero-Shot 3D Visual Grounding from Vision-Language Models
von: Li, Rong, et al.
Veröffentlicht: (2025)
von: Li, Rong, et al.
Veröffentlicht: (2025)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Visual Contextual Adaptation
von: Yu, Ming-Ming, et al.
Veröffentlicht: (2025)
von: Yu, Ming-Ming, et al.
Veröffentlicht: (2025)
SASI: Leveraging Sub-Action Semantics for Robust Early Action Recognition in Human-Robot Interaction
von: Cao, Yongpeng, et al.
Veröffentlicht: (2026)
von: Cao, Yongpeng, et al.
Veröffentlicht: (2026)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
von: Ao, Yuzhuo, et al.
Veröffentlicht: (2026)
von: Ao, Yuzhuo, et al.
Veröffentlicht: (2026)
Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
von: Qi, Xiuxiu, et al.
Veröffentlicht: (2025)
von: Qi, Xiuxiu, et al.
Veröffentlicht: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation
von: Jiang, Zhaodong, et al.
Veröffentlicht: (2025)
von: Jiang, Zhaodong, et al.
Veröffentlicht: (2025)
ZeroSCD: Zero-Shot Street Scene Change Detection
von: Kannan, Shyam Sundar, et al.
Veröffentlicht: (2024)
von: Kannan, Shyam Sundar, et al.
Veröffentlicht: (2024)
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
von: Zheng, Wanrong, et al.
Veröffentlicht: (2026)
von: Zheng, Wanrong, et al.
Veröffentlicht: (2026)
Leveraging YOLO-World and GPT-4V LMMs for Zero-Shot Person Detection and Action Recognition in Drone Imagery
von: Limberg, Christian, et al.
Veröffentlicht: (2024)
von: Limberg, Christian, et al.
Veröffentlicht: (2024)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs
von: Orjuela, Daniel Yezid Guarnizo, et al.
Veröffentlicht: (2026)
von: Orjuela, Daniel Yezid Guarnizo, et al.
Veröffentlicht: (2026)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping
von: Iwase, Shun, et al.
Veröffentlicht: (2025)
von: Iwase, Shun, et al.
Veröffentlicht: (2025)
EagleVision: A Multi-Task Benchmark for Cross-Domain Perception in High-Speed Autonomous Racing
von: Yagudin, Zakhar, et al.
Veröffentlicht: (2026)
von: Yagudin, Zakhar, et al.
Veröffentlicht: (2026)
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
von: Lyu, Kailin, et al.
Veröffentlicht: (2026)
von: Lyu, Kailin, et al.
Veröffentlicht: (2026)
Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models
von: Khalil, Aws, et al.
Veröffentlicht: (2026)
von: Khalil, Aws, et al.
Veröffentlicht: (2026)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
HARP-NeXt: High-Speed and Accurate Range-Point Fusion Network for 3D LiDAR Semantic Segmentation
von: Haidar, Samir Abou, et al.
Veröffentlicht: (2025)
von: Haidar, Samir Abou, et al.
Veröffentlicht: (2025)
NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion
von: Gregorek, Jakub, et al.
Veröffentlicht: (2026) -
ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control
von: Wang, Haosen, et al.
Veröffentlicht: (2026) -
MLFM: Multi-Layered Feature Maps for Richer Language Understanding in Zero-Shot Semantic Navigation
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2025) -
DriveVA: Video Action Models are Zero-Shot Drivers
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026) -
Improving Zero-Shot ObjectNav with Generative Communication
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2024)