Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Pani, Anupam, Yang, Yanchao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
di: Pani, Anupam, et al.
Pubblicazione: (2026)
di: Pani, Anupam, et al.
Pubblicazione: (2026)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
di: Pani, Anupam, et al.
Pubblicazione: (2025)
di: Pani, Anupam, et al.
Pubblicazione: (2025)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
di: Liu, Jiaming, et al.
Pubblicazione: (2024)
di: Liu, Jiaming, et al.
Pubblicazione: (2024)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
di: Din, Muhayy Ud, et al.
Pubblicazione: (2025)
di: Din, Muhayy Ud, et al.
Pubblicazione: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
di: Wang, Shijing, et al.
Pubblicazione: (2025)
di: Wang, Shijing, et al.
Pubblicazione: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
di: Li, Runhao, et al.
Pubblicazione: (2025)
di: Li, Runhao, et al.
Pubblicazione: (2025)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
di: Shao, Rui, et al.
Pubblicazione: (2025)
di: Shao, Rui, et al.
Pubblicazione: (2025)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
di: Chopra, Samarth, et al.
Pubblicazione: (2025)
di: Chopra, Samarth, et al.
Pubblicazione: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
di: Dai, Tingjun, et al.
Pubblicazione: (2026)
di: Dai, Tingjun, et al.
Pubblicazione: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
di: Wang, Guokang, et al.
Pubblicazione: (2024)
di: Wang, Guokang, et al.
Pubblicazione: (2024)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
di: Wang, Hengfei, et al.
Pubblicazione: (2026)
di: Wang, Hengfei, et al.
Pubblicazione: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
di: Wang, Shijing, et al.
Pubblicazione: (2026)
di: Wang, Shijing, et al.
Pubblicazione: (2026)
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
di: Zhou, Hanyu, et al.
Pubblicazione: (2025)
di: Zhou, Hanyu, et al.
Pubblicazione: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
di: Li, Chenxuan, et al.
Pubblicazione: (2024)
di: Li, Chenxuan, et al.
Pubblicazione: (2024)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
di: Mathew, Athul M., et al.
Pubblicazione: (2025)
di: Mathew, Athul M., et al.
Pubblicazione: (2025)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
di: Song, Zijian, et al.
Pubblicazione: (2025)
di: Song, Zijian, et al.
Pubblicazione: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
di: Xie, Haozhe, et al.
Pubblicazione: (2026)
di: Xie, Haozhe, et al.
Pubblicazione: (2026)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
di: Li, Qixiu, et al.
Pubblicazione: (2024)
di: Li, Qixiu, et al.
Pubblicazione: (2024)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
di: Yang, Shuai, et al.
Pubblicazione: (2025)
di: Yang, Shuai, et al.
Pubblicazione: (2025)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
di: Cheng, Hao, et al.
Pubblicazione: (2024)
di: Cheng, Hao, et al.
Pubblicazione: (2024)
VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
di: Bian, Jinyue, et al.
Pubblicazione: (2025)
di: Bian, Jinyue, et al.
Pubblicazione: (2025)
Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following
di: Gupta, Anshul, et al.
Pubblicazione: (2024)
di: Gupta, Anshul, et al.
Pubblicazione: (2024)
What Matters in Building Vision-Language-Action Models for Generalist Robots
di: Li, Xinghang, et al.
Pubblicazione: (2024)
di: Li, Xinghang, et al.
Pubblicazione: (2024)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
di: Zhao, Ruiteng, et al.
Pubblicazione: (2026)
di: Zhao, Ruiteng, et al.
Pubblicazione: (2026)
Physically Grounded Vision-Language Models for Robotic Manipulation
di: Gao, Jensen, et al.
Pubblicazione: (2023)
di: Gao, Jensen, et al.
Pubblicazione: (2023)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
di: Duan, Jiafei, et al.
Pubblicazione: (2024)
di: Duan, Jiafei, et al.
Pubblicazione: (2024)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze Information and Motion Bottlenecks
di: Takizawa, Ryo, et al.
Pubblicazione: (2025)
di: Takizawa, Ryo, et al.
Pubblicazione: (2025)
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
di: Guo, Shanshan, et al.
Pubblicazione: (2025)
di: Guo, Shanshan, et al.
Pubblicazione: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
di: Pani, Anupam, et al.
Pubblicazione: (2026) -
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
di: Pani, Anupam, et al.
Pubblicazione: (2025) -
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
di: Liu, Jiaming, et al.
Pubblicazione: (2024) -
Vision Language Action Models in Robotic Manipulation: A Systematic Review
di: Din, Muhayy Ud, et al.
Pubblicazione: (2025) -
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
di: Wang, Hongyu, et al.
Pubblicazione: (2025)