Benchmarking Cross-Domain Audio-Visual Deception Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Xiaobao, Yu, Zitong, Selvaraj, Nithish Muthuchamy, Shen, Bingquan, Kong, Adams Wai-Kin, Kot, Alex C.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909702818889728
author Guo, Xiaobao
Yu, Zitong
Selvaraj, Nithish Muthuchamy
Shen, Bingquan
Kong, Adams Wai-Kin
Kot, Alex C.
author_facet Guo, Xiaobao
Yu, Zitong
Selvaraj, Nithish Muthuchamy
Shen, Bingquan
Kong, Adams Wai-Kin
Kot, Alex C.
contents Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine the authenticity of an individual's statements. Nevertheless, recent developments in automated deception detection have demonstrated that multimodal features derived from both audio and video modalities may outperform human observers on publicly available datasets. Despite these positive findings, the generalizability of existing audio-visual deception detection approaches across different scenarios remains largely unexplored. To close this gap, we present the first cross-domain audio-visual deception detection benchmark, that enables us to assess how well these methods generalize for use in real-world scenarios. We used widely adopted audio and visual features and different architectures for benchmarking, comparing single-to-single and multi-to-single domain generalization performance. To further exploit the impacts using data from multiple source domains for training, we investigate three types of domain sampling strategies, including domain-simultaneous, domain-alternating, and domain-by-domain for multi-to-single domain generalization evaluation. We also propose an algorithm to enhance the generalization performance by maximizing the gradient inner products between modality encoders, named ``MM-IDGM". Furthermore, we proposed the Attention-Mixer fusion method to improve performance, and we believe that this new cross-domain benchmark will facilitate future research in audio-visual deception detection.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06995
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking Cross-Domain Audio-Visual Deception Detection
Guo, Xiaobao
Yu, Zitong
Selvaraj, Nithish Muthuchamy
Shen, Bingquan
Kong, Adams Wai-Kin
Kot, Alex C.
Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine the authenticity of an individual's statements. Nevertheless, recent developments in automated deception detection have demonstrated that multimodal features derived from both audio and video modalities may outperform human observers on publicly available datasets. Despite these positive findings, the generalizability of existing audio-visual deception detection approaches across different scenarios remains largely unexplored. To close this gap, we present the first cross-domain audio-visual deception detection benchmark, that enables us to assess how well these methods generalize for use in real-world scenarios. We used widely adopted audio and visual features and different architectures for benchmarking, comparing single-to-single and multi-to-single domain generalization performance. To further exploit the impacts using data from multiple source domains for training, we investigate three types of domain sampling strategies, including domain-simultaneous, domain-alternating, and domain-by-domain for multi-to-single domain generalization evaluation. We also propose an algorithm to enhance the generalization performance by maximizing the gradient inner products between modality encoders, named ``MM-IDGM". Furthermore, we proposed the Attention-Mixer fusion method to improve performance, and we believe that this new cross-domain benchmark will facilitate future research in audio-visual deception detection.
title Benchmarking Cross-Domain Audio-Visual Deception Detection
topic Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2405.06995