Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Andrew, Chuang, Ian, Gao, Dechen, Fukazawa, Kai, Soltani, Iman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
by: Chuang, Ian, et al.
Published: (2025)
by: Chuang, Ian, et al.
Published: (2025)
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments
by: Kazemi, Ehsan, et al.
Published: (2024)
by: Kazemi, Ehsan, et al.
Published: (2024)
From Scene to Object: Text-Guided Dual-Gaze Prediction
by: Ke, Zehong, et al.
Published: (2026)
by: Ke, Zehong, et al.
Published: (2026)
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
by: Tian, Yufeng, et al.
Published: (2026)
by: Tian, Yufeng, et al.
Published: (2026)
Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning
by: Shin, Ukcheol, et al.
Published: (2025)
by: Shin, Ukcheol, et al.
Published: (2025)
Automating Infrastructure Surveying: A Framework for Geometric Measurements and Compliance Assessment Using Point Cloud Data
by: Ghafourian, Amin, et al.
Published: (2025)
by: Ghafourian, Amin, et al.
Published: (2025)
Shape Completion and Real-Time Visualization in Robotic Ultrasound Spine Acquisitions
by: Gafencu, Miruna-Alexandra, et al.
Published: (2025)
by: Gafencu, Miruna-Alexandra, et al.
Published: (2025)
Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
by: Muttaqien, Muhammad A., et al.
Published: (2025)
by: Muttaqien, Muhammad A., et al.
Published: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
FloNa: Floor Plan Guided Embodied Visual Navigation
by: Li, Jiaxin, et al.
Published: (2024)
by: Li, Jiaxin, et al.
Published: (2024)
RDD4D: 4D Attention-Guided Road Damage Detection And Classification
by: Alkalbani, Asma, et al.
Published: (2025)
by: Alkalbani, Asma, et al.
Published: (2025)
Pair-VPR: Place-Aware Pre-training and Contrastive Pair Classification for Visual Place Recognition with Vision Transformers
by: Hausler, Stephen, et al.
Published: (2024)
by: Hausler, Stephen, et al.
Published: (2024)
EGSA-PT:Edge-Guided Spatial Attention with Progressive Training for Monocular Depth Estimation and Segmentation of Transparent Objects
by: Omotara, Gbenga, et al.
Published: (2025)
by: Omotara, Gbenga, et al.
Published: (2025)
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
by: Spigler, Giacomo
Published: (2026)
by: Spigler, Giacomo
Published: (2026)
RVN-Bench: A Benchmark for Reactive Visual Navigation
by: Lee, Jaewon, et al.
Published: (2026)
by: Lee, Jaewon, et al.
Published: (2026)
MCRL4OR: Multimodal Contrastive Representation Learning for Off-Road Environmental Perception
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Towards Deviation-Robust Agent Navigation via Perturbation-Aware Contrastive Learning
by: Lin, Bingqian, et al.
Published: (2024)
by: Lin, Bingqian, et al.
Published: (2024)
GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation
by: Wu, Ruihai, et al.
Published: (2025)
by: Wu, Ruihai, et al.
Published: (2025)
LaB-CL: Localized and Balanced Contrastive Learning for improving parking slot detection
by: Jeong, U Jin, et al.
Published: (2024)
by: Jeong, U Jin, et al.
Published: (2024)
Prior Availability in Industrial Visual Sim-to-Real: A Review of CAD-Guided and CAD-Unavailable Regimes
by: Tao, Chenxi, et al.
Published: (2026)
by: Tao, Chenxi, et al.
Published: (2026)
Negative Prototypes Guided Contrastive Learning for WSOD
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
S$^3$M-Net: Joint Learning of Semantic Segmentation and Stereo Matching for Autonomous Driving
by: Wu, Zhiyuan, et al.
Published: (2024)
by: Wu, Zhiyuan, et al.
Published: (2024)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
by: Cui, Wenbo, et al.
Published: (2025)
by: Cui, Wenbo, et al.
Published: (2025)
Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
by: Yuan, Zhecheng, et al.
Published: (2024)
by: Yuan, Zhecheng, et al.
Published: (2024)
Instruction-Guided Visual Masking
by: Zheng, Jinliang, et al.
Published: (2024)
by: Zheng, Jinliang, et al.
Published: (2024)
ShapeICP: Iterative Category-level Object Pose and Shape Estimation from Depth
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
DenseMTL: Cross-task Attention Mechanism for Dense Multi-task Learning
by: Lopes, Ivan, et al.
Published: (2022)
by: Lopes, Ivan, et al.
Published: (2022)
Playing to Vision Foundation Model's Strengths in Stereo Matching
by: Liu, Chuang-Wei, et al.
Published: (2024)
by: Liu, Chuang-Wei, et al.
Published: (2024)
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
by: Chen, Zezhou, et al.
Published: (2025)
by: Chen, Zezhou, et al.
Published: (2025)
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields
by: Yang, Zhaoyang, et al.
Published: (2026)
by: Yang, Zhaoyang, et al.
Published: (2026)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
CarDreamer: Open-Source Learning Platform for World Model based Autonomous Driving
by: Gao, Dechen, et al.
Published: (2024)
by: Gao, Dechen, et al.
Published: (2024)
DNAct: Diffusion Guided Multi-Task 3D Policy Learning
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent
by: Yu, Che Rin, et al.
Published: (2025)
by: Yu, Che Rin, et al.
Published: (2025)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
by: Tian, Ran, et al.
Published: (2023)
by: Tian, Ran, et al.
Published: (2023)
Similar Items
-
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
by: Chuang, Ian, et al.
Published: (2025) -
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation
by: Lee, Andrew, et al.
Published: (2024) -
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025) -
MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments
by: Kazemi, Ehsan, et al.
Published: (2024) -
From Scene to Object: Text-Guided Dual-Gaze Prediction
by: Ke, Zehong, et al.
Published: (2026)