Deep Learning for Visual Speech Analysis: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheng, Changchong, Kuang, Gangyao, Bai, Liang, Hou, Chenping, Guo, Yulan, Xu, Xin, Pietikäinen, Matti, Liu, Li
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910366208884736
author Sheng, Changchong
Kuang, Gangyao
Bai, Liang
Hou, Chenping
Guo, Yulan
Xu, Xin
Pietikäinen, Matti
Liu, Li
author_facet Sheng, Changchong
Kuang, Gangyao
Bai, Liang
Hou, Chenping
Guo, Yulan
Xu, Xin
Pietikäinen, Matti
Liu, Li
contents Visual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment. As a powerful AI strategy, deep learning techniques have extensively promoted the development of visual speech learning. Over the past five years, numerous deep learning based methods have been proposed to address various problems in this area, especially automatic visual speech recognition and generation. To push forward future research on visual speech, this paper aims to present a comprehensive review of recent progress in deep learning methods on visual speech analysis. We cover different aspects of visual speech, including fundamental problems, challenges, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. Besides, we also identify gaps in current research and discuss inspiring future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2205_10839
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Deep Learning for Visual Speech Analysis: A Survey
Sheng, Changchong
Kuang, Gangyao
Bai, Liang
Hou, Chenping
Guo, Yulan
Xu, Xin
Pietikäinen, Matti
Liu, Li
Computer Vision and Pattern Recognition
Visual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment. As a powerful AI strategy, deep learning techniques have extensively promoted the development of visual speech learning. Over the past five years, numerous deep learning based methods have been proposed to address various problems in this area, especially automatic visual speech recognition and generation. To push forward future research on visual speech, this paper aims to present a comprehensive review of recent progress in deep learning methods on visual speech analysis. We cover different aspects of visual speech, including fundamental problems, challenges, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. Besides, we also identify gaps in current research and discuss inspiring future research directions.
title Deep Learning for Visual Speech Analysis: A Survey
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2205.10839