Multi-view Gaze Target Estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Miao, Qiaomu, Golani, Vivek Raju, Xu, Jingyi, Dutta, Progga Paromita, Hoai, Minh, Samaras, Dimitris
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918118702448640
author Miao, Qiaomu
Golani, Vivek Raju
Xu, Jingyi
Dutta, Progga Paromita
Hoai, Minh
Samaras, Dimitris
author_facet Miao, Qiaomu
Golani, Vivek Raju
Xu, Jingyi
Dutta, Progga Paromita
Hoai, Minh
Samaras, Dimitris
contents This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to improve accuracy and expand applicability, addressing limitations in existing single-view methods that face challenges such as face occlusion, target ambiguity, and out-of-view targets. Our method processes a pair of camera views as input, incorporating a Head Information Aggregation (HIA) module for leveraging head information from both views for more accurate gaze estimation, an Uncertainty-based Gaze Selection (UGS) for identifying the most reliable gaze output, and an Epipolar-based Scene Attention (ESA) module for cross-view background information sharing. This approach significantly outperforms single-view baselines, especially when the second camera provides a clear view of the person's face. Additionally, our method can estimate the gaze target in the first view using the image of the person in the second view only, a capability not possessed by single-view GTE methods. Furthermore, the paper introduces a multi-view dataset for developing and evaluating multi-view GTE methods. Data and code are available at https://www3.cs.stonybrook.edu/~cvl/multiview_gte.html
format Preprint
id arxiv_https___arxiv_org_abs_2508_05857
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-view Gaze Target Estimation
Miao, Qiaomu
Golani, Vivek Raju
Xu, Jingyi
Dutta, Progga Paromita
Hoai, Minh
Samaras, Dimitris
Computer Vision and Pattern Recognition
This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to improve accuracy and expand applicability, addressing limitations in existing single-view methods that face challenges such as face occlusion, target ambiguity, and out-of-view targets. Our method processes a pair of camera views as input, incorporating a Head Information Aggregation (HIA) module for leveraging head information from both views for more accurate gaze estimation, an Uncertainty-based Gaze Selection (UGS) for identifying the most reliable gaze output, and an Epipolar-based Scene Attention (ESA) module for cross-view background information sharing. This approach significantly outperforms single-view baselines, especially when the second camera provides a clear view of the person's face. Additionally, our method can estimate the gaze target in the first view using the image of the person in the second view only, a capability not possessed by single-view GTE methods. Furthermore, the paper introduces a multi-view dataset for developing and evaluating multi-view GTE methods. Data and code are available at https://www3.cs.stonybrook.edu/~cvl/multiview_gte.html
title Multi-view Gaze Target Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.05857