Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Pani, Anupam, Yang, Yanchao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
por: Pani, Anupam, et al.
Publicado: (2025)
por: Pani, Anupam, et al.
Publicado: (2025)
Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
por: Pani, Anupam, et al.
Publicado: (2026)
por: Pani, Anupam, et al.
Publicado: (2026)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
por: Fu, Yuqian, et al.
Publicado: (2024)
por: Fu, Yuqian, et al.
Publicado: (2024)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
por: Özsoy, Ege, et al.
Publicado: (2025)
por: Özsoy, Ege, et al.
Publicado: (2025)
Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
por: Huang, Junsheng, et al.
Publicado: (2025)
por: Huang, Junsheng, et al.
Publicado: (2025)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
por: Reilly, Dominick, et al.
Publicado: (2025)
por: Reilly, Dominick, et al.
Publicado: (2025)
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
por: Chen, Qinyu, et al.
Publicado: (2025)
por: Chen, Qinyu, et al.
Publicado: (2025)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
por: Zhu, Bingwen, et al.
Publicado: (2026)
por: Zhu, Bingwen, et al.
Publicado: (2026)
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
por: Ge, Mengmeng, et al.
Publicado: (2026)
por: Ge, Mengmeng, et al.
Publicado: (2026)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
por: Schäfer, Finn Rasmus, et al.
Publicado: (2026)
por: Schäfer, Finn Rasmus, et al.
Publicado: (2026)
EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
por: John, Ronan, et al.
Publicado: (2025)
por: John, Ronan, et al.
Publicado: (2025)
ECHO: Ego-Centric modeling of Human-Object interactions
por: Petrov, Ilya A., et al.
Publicado: (2025)
por: Petrov, Ilya A., et al.
Publicado: (2025)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
por: Bandraupalli, Srihari, et al.
Publicado: (2025)
por: Bandraupalli, Srihari, et al.
Publicado: (2025)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
por: Mahdi, Mohammad, et al.
Publicado: (2026)
por: Mahdi, Mohammad, et al.
Publicado: (2026)
GazeNLQ @ Ego4D Natural Language Queries Challenge 2025
por: Lin, Wei-Cheng, et al.
Publicado: (2025)
por: Lin, Wei-Cheng, et al.
Publicado: (2025)
ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios
por: Ragusa, Francesco, et al.
Publicado: (2026)
por: Ragusa, Francesco, et al.
Publicado: (2026)
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
por: Sharma, Pavan Kumar, et al.
Publicado: (2023)
por: Sharma, Pavan Kumar, et al.
Publicado: (2023)
EgoFSD: Ego-Centric Fully Sparse Paradigm with Uncertainty Denoising and Iterative Refinement for Efficient End-to-End Self-Driving
por: Su, Haisheng, et al.
Publicado: (2024)
por: Su, Haisheng, et al.
Publicado: (2024)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
por: Zhang, Deheng, et al.
Publicado: (2025)
por: Zhang, Deheng, et al.
Publicado: (2025)
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
por: Wang, Zeyu, et al.
Publicado: (2026)
por: Wang, Zeyu, et al.
Publicado: (2026)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
por: Gholami, Mohsen, et al.
Publicado: (2025)
por: Gholami, Mohsen, et al.
Publicado: (2025)
EgoAVU: Egocentric Audio-Visual Understanding
por: Seth, Ashish, et al.
Publicado: (2026)
por: Seth, Ashish, et al.
Publicado: (2026)
Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction
por: Akbiyik, M. Eren, et al.
Publicado: (2023)
por: Akbiyik, M. Eren, et al.
Publicado: (2023)
Contour Errors: An Ego-Centric Metric for Reliable 3D Multi-Object Tracking
por: Kaul, Sharang, et al.
Publicado: (2025)
por: Kaul, Sharang, et al.
Publicado: (2025)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
por: Li, Yiwei, et al.
Publicado: (2026)
por: Li, Yiwei, et al.
Publicado: (2026)
Event-Free Moving Object Segmentation from Moving Ego Vehicle
por: Zhou, Zhuyun, et al.
Publicado: (2023)
por: Zhou, Zhuyun, et al.
Publicado: (2023)
On the Perception Bottleneck of VLMs for Chart Understanding
por: Liu, Junteng, et al.
Publicado: (2025)
por: Liu, Junteng, et al.
Publicado: (2025)
CIVET: Systematic Evaluation of Understanding in VLMs
por: Rizzoli, Massimo, et al.
Publicado: (2025)
por: Rizzoli, Massimo, et al.
Publicado: (2025)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
por: Mahdi, Mohammad, et al.
Publicado: (2025)
por: Mahdi, Mohammad, et al.
Publicado: (2025)
Linear Scaling Video VLMs for Long Video Understanding
por: Eyzaguirre, Cristobal, et al.
Publicado: (2026)
por: Eyzaguirre, Cristobal, et al.
Publicado: (2026)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
por: Sun, Shitong, et al.
Publicado: (2026)
por: Sun, Shitong, et al.
Publicado: (2026)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
por: Zheng, Henry, et al.
Publicado: (2025)
por: Zheng, Henry, et al.
Publicado: (2025)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
por: Dai, Yang, et al.
Publicado: (2026)
por: Dai, Yang, et al.
Publicado: (2026)
Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction
por: Wei, Dongxu, et al.
Publicado: (2024)
por: Wei, Dongxu, et al.
Publicado: (2024)
EgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation
por: Zhang, Ganlin, et al.
Publicado: (2023)
por: Zhang, Ganlin, et al.
Publicado: (2023)
Improving Domain Generalization on Gaze Estimation via Branch-out Auxiliary Regularization
por: Zhao, Ruijie, et al.
Publicado: (2024)
por: Zhao, Ruijie, et al.
Publicado: (2024)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
por: Leonardi, Rosario, et al.
Publicado: (2026)
por: Leonardi, Rosario, et al.
Publicado: (2026)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
por: Fu, Yuqian, et al.
Publicado: (2025)
por: Fu, Yuqian, et al.
Publicado: (2025)
Understanding and Modeling the Effects of Task and Context on Drivers' Gaze Allocation
por: Kotseruba, Iuliia, et al.
Publicado: (2023)
por: Kotseruba, Iuliia, et al.
Publicado: (2023)
MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
por: Dai, Ming, et al.
Publicado: (2025)
por: Dai, Ming, et al.
Publicado: (2025)
Ejemplares similares
-
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
por: Pani, Anupam, et al.
Publicado: (2025) -
Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
por: Pani, Anupam, et al.
Publicado: (2026) -
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
por: Fu, Yuqian, et al.
Publicado: (2024) -
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
por: Özsoy, Ege, et al.
Publicado: (2025) -
Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
por: Huang, Junsheng, et al.
Publicado: (2025)