A multi-modal approach for identifying schizophrenia using cross-modal attention

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Premananth, Gowtham, Siriwardena, Yashish M., Resnik, Philip, Espy-Wilson, Carol
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914760794046464
author Premananth, Gowtham
Siriwardena, Yashish M.
Resnik, Philip
Espy-Wilson, Carol
author_facet Premananth, Gowtham
Siriwardena, Yashish M.
Resnik, Philip
Espy-Wilson, Carol
contents This study focuses on how different modalities of human communication can be used to distinguish between healthy controls and subjects with schizophrenia who exhibit strong positive symptoms. We developed a multi-modal schizophrenia classification system using audio, video, and text. Facial action units and vocal tract variables were extracted as low-level features from video and audio respectively, which were then used to compute high-level coordination features that served as the inputs to the audio and video modalities. Context-independent text embeddings extracted from transcriptions of speech were used as the input for the text modality. The multi-modal system is developed by fusing a segment-to-session-level classifier for video and audio modalities with a text model based on a Hierarchical Attention Network (HAN) with cross-modal attention. The proposed multi-modal system outperforms the previous state-of-the-art multi-modal system by 8.53% in the weighted average F1 score.
format Preprint
id arxiv_https___arxiv_org_abs_2309_15136
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A multi-modal approach for identifying schizophrenia using cross-modal attention
Premananth, Gowtham
Siriwardena, Yashish M.
Resnik, Philip
Espy-Wilson, Carol
Signal Processing
Multimedia
Sound
Audio and Speech Processing
Image and Video Processing
This study focuses on how different modalities of human communication can be used to distinguish between healthy controls and subjects with schizophrenia who exhibit strong positive symptoms. We developed a multi-modal schizophrenia classification system using audio, video, and text. Facial action units and vocal tract variables were extracted as low-level features from video and audio respectively, which were then used to compute high-level coordination features that served as the inputs to the audio and video modalities. Context-independent text embeddings extracted from transcriptions of speech were used as the input for the text modality. The multi-modal system is developed by fusing a segment-to-session-level classifier for video and audio modalities with a text model based on a Hierarchical Attention Network (HAN) with cross-modal attention. The proposed multi-modal system outperforms the previous state-of-the-art multi-modal system by 8.53% in the weighted average F1 score.
title A multi-modal approach for identifying schizophrenia using cross-modal attention
topic Signal Processing
Multimedia
Sound
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2309.15136