Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Chuang, Ian, Zou, Jinyu, Lee, Andrew, Gao, Dechen, Soltani, Iman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation
by: Chuang, Ian, et al.
Published: (2024)
by: Chuang, Ian, et al.
Published: (2024)
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments
by: Kazemi, Ehsan, et al.
Published: (2024)
by: Kazemi, Ehsan, et al.
Published: (2024)
Automating Infrastructure Surveying: A Framework for Geometric Measurements and Compliance Assessment Using Point Cloud Data
by: Ghafourian, Amin, et al.
Published: (2025)
by: Ghafourian, Amin, et al.
Published: (2025)
Reading in the Dark with Foveated Event Vision
by: Brander, Carl, et al.
Published: (2025)
by: Brander, Carl, et al.
Published: (2025)
Low-Cost Infrared Vision Systems for Improved Safety of Emergency Vehicle Operations Under Low-Visibility Conditions
by: Naddaf-Sh, M-Mahdi, et al.
Published: (2025)
by: Naddaf-Sh, M-Mahdi, et al.
Published: (2025)
CarDreamer: Open-Source Learning Platform for World Model based Autonomous Driving
by: Gao, Dechen, et al.
Published: (2024)
by: Gao, Dechen, et al.
Published: (2024)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025)
by: Kerr, Justin, et al.
Published: (2025)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
by: Zuo, Kuangji, et al.
Published: (2026)
by: Zuo, Kuangji, et al.
Published: (2026)
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze Information and Motion Bottlenecks
by: Takizawa, Ryo, et al.
Published: (2025)
by: Takizawa, Ryo, et al.
Published: (2025)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024)
by: Wang, Guokang, et al.
Published: (2024)
IN-RIL: Interleaved Reinforcement and Imitation Learning for Policy Fine-Tuning
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
Learning Priors of Human Motion With Vision Transformers
by: Falqueto, Placido, et al.
Published: (2025)
by: Falqueto, Placido, et al.
Published: (2025)
ROI-Driven Foveated Attention for Unified Egocentric Representations in Vision-Language-Action Systems
by: Sun, Xinhai, et al.
Published: (2026)
by: Sun, Xinhai, et al.
Published: (2026)
TesserAct: Learning 4D Embodied World Models
by: Zhen, Haoyu, et al.
Published: (2025)
by: Zhen, Haoyu, et al.
Published: (2025)
Gaze Estimation for Human-Robot Interaction: Analysis Using the NICO Platform
by: Palider, Matej, et al.
Published: (2025)
by: Palider, Matej, et al.
Published: (2025)
Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction
by: Sezer, Berk, et al.
Published: (2026)
by: Sezer, Berk, et al.
Published: (2026)
GazeVLA: Learning Human Intention for Robotic Manipulation
by: Li, Chengyang, et al.
Published: (2026)
by: Li, Chengyang, et al.
Published: (2026)
Humanizing Robot Gaze Shifts: A Framework for Natural Gaze Shifts in Humanoid Robots
by: Wei, Jingchao, et al.
Published: (2026)
by: Wei, Jingchao, et al.
Published: (2026)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
by: Bai, Yongjie, et al.
Published: (2025)
by: Bai, Yongjie, et al.
Published: (2025)
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025)
by: Yao, Yuanqi, et al.
Published: (2025)
Where Do We Look When We Teach? Analyzing Human Gaze Behavior Across Demonstration Devices in Robot Imitation Learning
by: Ishida, Yutaro, et al.
Published: (2025)
by: Ishida, Yutaro, et al.
Published: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
by: Xue, Wei, et al.
Published: (2026)
by: Xue, Wei, et al.
Published: (2026)
Gaze Detection and Analysis for Initiating Joint Activity in Industrial Human-Robot Collaboration
by: Prajod, Pooja, et al.
Published: (2023)
by: Prajod, Pooja, et al.
Published: (2023)
ViTA-Seg: Vision Transformer for Amodal Segmentation in Robotics
by: Caramia, Donato, et al.
Published: (2025)
by: Caramia, Donato, et al.
Published: (2025)
Vision-Based Safe Human-Robot Collaboration with Uncertainty Guarantees
by: Thumm, Jakob, et al.
Published: (2026)
by: Thumm, Jakob, et al.
Published: (2026)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2026)
by: Huang, Chi-Pin, et al.
Published: (2026)
Bidirectional Human-Robot Communication for Physical Human-Robot Interaction
by: Wang, Junxiang, et al.
Published: (2026)
by: Wang, Junxiang, et al.
Published: (2026)
Krysalis Hand: A Lightweight, High-Payload, 18-DoF Anthropomorphic End-Effector for Robotic Learning and Dexterous Manipulation
by: Basheer, Al Arsh, et al.
Published: (2025)
by: Basheer, Al Arsh, et al.
Published: (2025)
Symmetry-Aware Fusion of Vision and Tactile Sensing via Bilateral Force Priors for Robotic Manipulation
by: Lee, Wonju, et al.
Published: (2026)
by: Lee, Wonju, et al.
Published: (2026)
GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering
by: Ebadulla, Farhaan, et al.
Published: (2025)
by: Ebadulla, Farhaan, et al.
Published: (2025)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
by: Yang, Jiashu, et al.
Published: (2025)
by: Yang, Jiashu, et al.
Published: (2025)
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
by: Gong, Litian, et al.
Published: (2025)
by: Gong, Litian, et al.
Published: (2025)
HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction
by: Shi, Zhonghao, et al.
Published: (2025)
by: Shi, Zhonghao, et al.
Published: (2025)
Adversarial Attacks and Detection in Visual Place Recognition for Safer Robot Navigation
by: Malone, Connor, et al.
Published: (2025)
by: Malone, Connor, et al.
Published: (2025)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
by: Ma, Teli, et al.
Published: (2025)
by: Ma, Teli, et al.
Published: (2025)
Gaze Behavior During a Long-Term, In-Home, Social Robot Intervention for Children with ASD
by: Ramnauth, Rebecca, et al.
Published: (2025)
by: Ramnauth, Rebecca, et al.
Published: (2025)
Similar Items
-
Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning
by: Lee, Andrew, et al.
Published: (2025) -
Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation
by: Chuang, Ian, et al.
Published: (2024) -
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025) -
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation
by: Lee, Andrew, et al.
Published: (2024) -
MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments
by: Kazemi, Ehsan, et al.
Published: (2024)