TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
Fuente:
arXiv
Saved in:
| Main Author: | Spigler, Giacomo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predicting Depression and Anxiety Risk in Dutch Neighborhoods from Street-View Images
by: Khodorivsko, Nin, et al.
Published: (2024)
by: Khodorivsko, Nin, et al.
Published: (2024)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
EmbodiSwap for Zero-Shot Robot Imitation Learning
by: Dessalene, Eadom, et al.
Published: (2025)
by: Dessalene, Eadom, et al.
Published: (2025)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
by: Tomilin, Tristan, et al.
Published: (2025)
by: Tomilin, Tristan, et al.
Published: (2025)
Instant Policy: In-Context Imitation Learning via Graph Diffusion
by: Vosylius, Vitalis, et al.
Published: (2024)
by: Vosylius, Vitalis, et al.
Published: (2024)
What's the Move? Hybrid Imitation Learning via Salient Points
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
M2Distill: Multi-Modal Distillation for Lifelong Imitation Learning
by: Roy, Kaushik, et al.
Published: (2024)
by: Roy, Kaushik, et al.
Published: (2024)
SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation
by: Zhou, Zihan, et al.
Published: (2024)
by: Zhou, Zihan, et al.
Published: (2024)
NIL: No-data Imitation Learning by Leveraging Pre-trained Video Diffusion Models
by: Albaba, Mert, et al.
Published: (2025)
by: Albaba, Mert, et al.
Published: (2025)
Imitation of human motion achieves natural head movements for humanoid robots in an active-speaker detection task
by: Ding, Bosong, et al.
Published: (2024)
by: Ding, Bosong, et al.
Published: (2024)
3D Hand Pose Estimation in Everyday Egocentric Images
by: Prakash, Aditya, et al.
Published: (2023)
by: Prakash, Aditya, et al.
Published: (2023)
DeFIX: Detecting and Fixing Failure Scenarios with Reinforcement Learning in Imitation Learning Based Autonomous Driving
by: Dagdanov, Resul, et al.
Published: (2022)
by: Dagdanov, Resul, et al.
Published: (2022)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
by: Patel, Alkesh, et al.
Published: (2025)
by: Patel, Alkesh, et al.
Published: (2025)
DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning
by: Jiang, Zhenyu, et al.
Published: (2024)
by: Jiang, Zhenyu, et al.
Published: (2024)
Adapt3R: Adaptive 3D Scene Representation for Domain Transfer in Imitation Learning
by: Wilcox, Albert, et al.
Published: (2025)
by: Wilcox, Albert, et al.
Published: (2025)
What Matters to Enhance Traffic Rule Compliance of Imitation Learning for End-to-End Autonomous Driving
by: Zhou, Hongkuan, et al.
Published: (2023)
by: Zhou, Hongkuan, et al.
Published: (2023)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
by: Goyal, Divyanshu, et al.
Published: (2026)
by: Goyal, Divyanshu, et al.
Published: (2026)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Observer-Actor: Active Vision Imitation Learning with Sparse-View Gaussian Splatting
by: Wang, Yilong, et al.
Published: (2025)
by: Wang, Yilong, et al.
Published: (2025)
OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
by: Li, Jinhan, et al.
Published: (2024)
by: Li, Jinhan, et al.
Published: (2024)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
by: Park, Yohan, et al.
Published: (2025)
by: Park, Yohan, et al.
Published: (2025)
SegXAL: Explainable Active Learning for Semantic Segmentation in Driving Scene Scenarios
by: Mandalika, Sriram, et al.
Published: (2024)
by: Mandalika, Sriram, et al.
Published: (2024)
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
Theia: Distilling Diverse Vision Foundation Models for Robot Learning
by: Shang, Jinghuan, et al.
Published: (2024)
by: Shang, Jinghuan, et al.
Published: (2024)
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
by: Grotz, Markus, et al.
Published: (2024)
by: Grotz, Markus, et al.
Published: (2024)
EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
by: Punamiya, Ryan, et al.
Published: (2025)
by: Punamiya, Ryan, et al.
Published: (2025)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation
by: Kim, Minwoo, et al.
Published: (2025)
by: Kim, Minwoo, et al.
Published: (2025)
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
by: Athalye, Ashay, et al.
Published: (2024)
by: Athalye, Ashay, et al.
Published: (2024)
AI Guide Dog: Egocentric Path Prediction on Smartphone
by: Jadhav, Aishwarya, et al.
Published: (2025)
by: Jadhav, Aishwarya, et al.
Published: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
SurgPose: Generalisable Surgical Instrument Pose Estimation using Zero-Shot Learning and Stereo Vision
by: Rai, Utsav, et al.
Published: (2025)
by: Rai, Utsav, et al.
Published: (2025)
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
by: Vasa, Santosh, et al.
Published: (2025)
by: Vasa, Santosh, et al.
Published: (2025)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
KINESIS: Motion Imitation for Human Musculoskeletal Locomotion
by: Simos, Merkourios, et al.
Published: (2025)
by: Simos, Merkourios, et al.
Published: (2025)
Extrapolated Urban View Synthesis Benchmark
by: Han, Xiangyu, et al.
Published: (2024)
by: Han, Xiangyu, et al.
Published: (2024)
A Vision-Enabled Prosthetic Hand for Children with Upper Limb Disabilities
by: Sarker, Md Abdul Baset, et al.
Published: (2025)
by: Sarker, Md Abdul Baset, et al.
Published: (2025)
Similar Items
-
Predicting Depression and Anxiety Risk in Dutch Neighborhoods from Street-View Images
by: Khodorivsko, Nin, et al.
Published: (2024) -
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
by: Yang, Ruihan, et al.
Published: (2025) -
EmbodiSwap for Zero-Shot Robot Imitation Learning
by: Dessalene, Eadom, et al.
Published: (2025) -
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
by: Tomilin, Tristan, et al.
Published: (2025) -
Instant Policy: In-Context Imitation Learning via Graph Diffusion
by: Vosylius, Vitalis, et al.
Published: (2024)