Saved in:
Bibliographic Details
Main Authors: Lubitz, Adrian, Valdenegro-Toro, Matias, Kirchner, Frank
Format: Preprint
Published: 2021
Subjects:
Online Access:https://arxiv.org/abs/2109.13789
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929195973607424
author Lubitz, Adrian
Valdenegro-Toro, Matias
Kirchner, Frank
author_facet Lubitz, Adrian
Valdenegro-Toro, Matias
Kirchner, Frank
contents Robots are becoming everyday devices, increasing their interaction with humans. To make human-machine interaction more natural, cognitive features like Visual Voice Activity Detection (VVAD), which can detect whether a person is speaking or not, given visual input of a camera, need to be implemented. Neural networks are state of the art for tasks in Image Processing, Time Series Prediction, Natural Language Processing and other domains. Those Networks require large quantities of labeled data. Currently there are not many datasets for the task of VVAD. In this work we created a large scale dataset called the VVAD-LRS3 dataset, derived by automatic annotations from the LRS3 dataset. The VVAD-LRS3 dataset contains over 44K samples, over three times the next competitive dataset (WildVVAD). We evaluate different baselines on four kinds of features: facial and lip images, and facial and lip landmark features. With a Convolutional Neural Network Long Short Term Memory (CNN LSTM) on facial images an accuracy of 92% was reached on the test set. A study with humans showed that they reach an accuracy of 87.93% on the test set.
format Preprint
id arxiv_https___arxiv_org_abs_2109_13789
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle The VVAD-LRS3 Dataset for Visual Voice Activity Detection
Lubitz, Adrian
Valdenegro-Toro, Matias
Kirchner, Frank
Computer Vision and Pattern Recognition
Robotics
Robots are becoming everyday devices, increasing their interaction with humans. To make human-machine interaction more natural, cognitive features like Visual Voice Activity Detection (VVAD), which can detect whether a person is speaking or not, given visual input of a camera, need to be implemented. Neural networks are state of the art for tasks in Image Processing, Time Series Prediction, Natural Language Processing and other domains. Those Networks require large quantities of labeled data. Currently there are not many datasets for the task of VVAD. In this work we created a large scale dataset called the VVAD-LRS3 dataset, derived by automatic annotations from the LRS3 dataset. The VVAD-LRS3 dataset contains over 44K samples, over three times the next competitive dataset (WildVVAD). We evaluate different baselines on four kinds of features: facial and lip images, and facial and lip landmark features. With a Convolutional Neural Network Long Short Term Memory (CNN LSTM) on facial images an accuracy of 92% was reached on the test set. A study with humans showed that they reach an accuracy of 87.93% on the test set.
title The VVAD-LRS3 Dataset for Visual Voice Activity Detection
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2109.13789