Multi-Modal Gaze Following in Conversational Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Yuqi, Zhang, Zhongqun, Horanyi, Nora, Moon, Jaewon, Cheng, Yihua, Chang, Hyung Jin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909074695651328
author Hou, Yuqi
Zhang, Zhongqun
Horanyi, Nora
Moon, Jaewon
Cheng, Yihua
Chang, Hyung Jin
author_facet Hou, Yuqi
Zhang, Zhongqun
Horanyi, Nora
Moon, Jaewon
Cheng, Yihua
Chang, Hyung Jin
contents Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides crucial cues for determining human behavior.This suggests that we can further improve gaze following considering audio cues. In this paper, we explore gaze following tasks in conversational scenarios. We propose a novel multi-modal gaze following framework based on our observation ``audiences tend to focus on the speaker''. We first leverage the correlation between audio and lips, and classify speakers and listeners in a scene. We then use the identity information to enhance scene images and propose a gaze candidate estimation network. The network estimates gaze candidates from enhanced scene images and we use MLP to match subjects with candidates as classification tasks. Existing gaze following datasets focus on visual images while ignore audios.To evaluate our method, we collect a conversational dataset, VideoGazeSpeech (VGS), which is the first gaze following dataset including images and audio. Our method significantly outperforms existing methods in VGS datasets. The visualization result also prove the advantage of audio cues in gaze following tasks. Our work will inspire more researches in multi-modal gaze following estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2311_05669
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Multi-Modal Gaze Following in Conversational Scenarios
Hou, Yuqi
Zhang, Zhongqun
Horanyi, Nora
Moon, Jaewon
Cheng, Yihua
Chang, Hyung Jin
Computer Vision and Pattern Recognition
Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides crucial cues for determining human behavior.This suggests that we can further improve gaze following considering audio cues. In this paper, we explore gaze following tasks in conversational scenarios. We propose a novel multi-modal gaze following framework based on our observation ``audiences tend to focus on the speaker''. We first leverage the correlation between audio and lips, and classify speakers and listeners in a scene. We then use the identity information to enhance scene images and propose a gaze candidate estimation network. The network estimates gaze candidates from enhanced scene images and we use MLP to match subjects with candidates as classification tasks. Existing gaze following datasets focus on visual images while ignore audios.To evaluate our method, we collect a conversational dataset, VideoGazeSpeech (VGS), which is the first gaze following dataset including images and audio. Our method significantly outperforms existing methods in VGS datasets. The visualization result also prove the advantage of audio cues in gaze following tasks. Our work will inspire more researches in multi-modal gaze following estimation.
title Multi-Modal Gaze Following in Conversational Scenarios
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.05669