Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
Fuente:
arXiv
Guardado en:
| Autores principales: | Pani, Anupam, Yang, Yanchao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
por: Pani, Anupam, et al.
Publicado: (2026)
por: Pani, Anupam, et al.
Publicado: (2026)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
por: Pani, Anupam, et al.
Publicado: (2025)
por: Pani, Anupam, et al.
Publicado: (2025)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
por: Liu, Jiaming, et al.
Publicado: (2024)
por: Liu, Jiaming, et al.
Publicado: (2024)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
por: Din, Muhayy Ud, et al.
Publicado: (2025)
por: Din, Muhayy Ud, et al.
Publicado: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
por: Wang, Hongyu, et al.
Publicado: (2025)
por: Wang, Hongyu, et al.
Publicado: (2025)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
por: Wang, Shijing, et al.
Publicado: (2025)
por: Wang, Shijing, et al.
Publicado: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
por: Shi, Hao, et al.
Publicado: (2025)
por: Shi, Hao, et al.
Publicado: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
por: Li, Runhao, et al.
Publicado: (2025)
por: Li, Runhao, et al.
Publicado: (2025)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
por: Shao, Rui, et al.
Publicado: (2025)
por: Shao, Rui, et al.
Publicado: (2025)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
por: Chopra, Samarth, et al.
Publicado: (2025)
por: Chopra, Samarth, et al.
Publicado: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
por: Kim, Ju-Young, et al.
Publicado: (2025)
por: Kim, Ju-Young, et al.
Publicado: (2025)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
por: Dai, Tingjun, et al.
Publicado: (2026)
por: Dai, Tingjun, et al.
Publicado: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
por: Liu, Mengzhen, et al.
Publicado: (2026)
por: Liu, Mengzhen, et al.
Publicado: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
por: Wen, Junjie, et al.
Publicado: (2024)
por: Wen, Junjie, et al.
Publicado: (2024)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
por: Wang, Guokang, et al.
Publicado: (2024)
por: Wang, Guokang, et al.
Publicado: (2024)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
por: Wang, Hengfei, et al.
Publicado: (2026)
por: Wang, Hengfei, et al.
Publicado: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
por: Ding, Pengxiang, et al.
Publicado: (2023)
por: Ding, Pengxiang, et al.
Publicado: (2023)
Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
por: Wang, Shijing, et al.
Publicado: (2026)
por: Wang, Shijing, et al.
Publicado: (2026)
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
por: Zhou, Hanyu, et al.
Publicado: (2025)
por: Zhou, Hanyu, et al.
Publicado: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
por: Li, Chenxuan, et al.
Publicado: (2024)
por: Li, Chenxuan, et al.
Publicado: (2024)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
por: Cheng, An-Chieh, et al.
Publicado: (2024)
por: Cheng, An-Chieh, et al.
Publicado: (2024)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
por: Mathew, Athul M., et al.
Publicado: (2025)
por: Mathew, Athul M., et al.
Publicado: (2025)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
por: Song, Zijian, et al.
Publicado: (2025)
por: Song, Zijian, et al.
Publicado: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
por: Li, Hao, et al.
Publicado: (2025)
por: Li, Hao, et al.
Publicado: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
por: Xie, Haozhe, et al.
Publicado: (2026)
por: Xie, Haozhe, et al.
Publicado: (2026)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
por: Li, Qixiu, et al.
Publicado: (2024)
por: Li, Qixiu, et al.
Publicado: (2024)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
por: Yang, Shuai, et al.
Publicado: (2025)
por: Yang, Shuai, et al.
Publicado: (2025)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
por: Cheng, Hao, et al.
Publicado: (2024)
por: Cheng, Hao, et al.
Publicado: (2024)
VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
por: Bian, Jinyue, et al.
Publicado: (2025)
por: Bian, Jinyue, et al.
Publicado: (2025)
Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following
por: Gupta, Anshul, et al.
Publicado: (2024)
por: Gupta, Anshul, et al.
Publicado: (2024)
What Matters in Building Vision-Language-Action Models for Generalist Robots
por: Li, Xinghang, et al.
Publicado: (2024)
por: Li, Xinghang, et al.
Publicado: (2024)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
por: Zhao, Ruiteng, et al.
Publicado: (2026)
por: Zhao, Ruiteng, et al.
Publicado: (2026)
Physically Grounded Vision-Language Models for Robotic Manipulation
por: Gao, Jensen, et al.
Publicado: (2023)
por: Gao, Jensen, et al.
Publicado: (2023)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
por: Duan, Jiafei, et al.
Publicado: (2024)
por: Duan, Jiafei, et al.
Publicado: (2024)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
por: Huang, Ting, et al.
Publicado: (2025)
por: Huang, Ting, et al.
Publicado: (2025)
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
por: Zhou, Jiaming, et al.
Publicado: (2025)
por: Zhou, Jiaming, et al.
Publicado: (2025)
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze Information and Motion Bottlenecks
por: Takizawa, Ryo, et al.
Publicado: (2025)
por: Takizawa, Ryo, et al.
Publicado: (2025)
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
por: Guo, Shanshan, et al.
Publicado: (2025)
por: Guo, Shanshan, et al.
Publicado: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
por: Song, Wenxuan, et al.
Publicado: (2025)
por: Song, Wenxuan, et al.
Publicado: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
por: Huang, Haifeng, et al.
Publicado: (2025)
por: Huang, Haifeng, et al.
Publicado: (2025)
Ejemplares similares
-
Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
por: Pani, Anupam, et al.
Publicado: (2026) -
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
por: Pani, Anupam, et al.
Publicado: (2025) -
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
por: Liu, Jiaming, et al.
Publicado: (2024) -
Vision Language Action Models in Robotic Manipulation: A Systematic Review
por: Din, Muhayy Ud, et al.
Publicado: (2025) -
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
por: Wang, Hongyu, et al.
Publicado: (2025)