MMFformer: Multimodal Fusion Transformer Network for Depression Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Haque, Md Rezwanul, Islam, Md. Milon, Raju, S M Taslim Uddin, Altaheri, Hamdi, Nassar, Lobna, Karray, Fakhri
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912532265959424
author Haque, Md Rezwanul
Islam, Md. Milon
Raju, S M Taslim Uddin
Altaheri, Hamdi
Nassar, Lobna
Karray, Fakhri
author_facet Haque, Md Rezwanul
Islam, Md. Milon
Raju, S M Taslim Uddin
Altaheri, Hamdi
Nassar, Lobna
Karray, Fakhri
contents Depression is a serious mental health illness that significantly affects an individual's well-being and quality of life, making early detection crucial for adequate care and treatment. Detecting depression is often difficult, as it is based primarily on subjective evaluations during clinical interviews. Hence, the early diagnosis of depression, thanks to the content of social networks, has become a prominent research area. The extensive and diverse nature of user-generated information poses a significant challenge, limiting the accurate extraction of relevant temporal information and the effective fusion of data across multiple modalities. This paper introduces MMFformer, a multimodal depression detection network designed to retrieve depressive spatio-temporal high-level patterns from multimodal social media information. The transformer network with residual connections captures spatial features from videos, and a transformer encoder is exploited to design important temporal dynamics in audio. Moreover, the fusion architecture fused the extracted features through late and intermediate fusion strategies to find out the most relevant intermodal correlations among them. Finally, the proposed network is assessed on two large-scale depression detection datasets, and the results clearly reveal that it surpasses existing state-of-the-art approaches, improving the F1-Score by 13.92% for D-Vlog dataset and 7.74% for LMVD dataset. The code is made available publicly at https://github.com/rezwanh001/Large-Scale-Multimodal-Depression-Detection.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06701
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMFformer: Multimodal Fusion Transformer Network for Depression Detection
Haque, Md Rezwanul
Islam, Md. Milon
Raju, S M Taslim Uddin
Altaheri, Hamdi
Nassar, Lobna
Karray, Fakhri
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
Depression is a serious mental health illness that significantly affects an individual's well-being and quality of life, making early detection crucial for adequate care and treatment. Detecting depression is often difficult, as it is based primarily on subjective evaluations during clinical interviews. Hence, the early diagnosis of depression, thanks to the content of social networks, has become a prominent research area. The extensive and diverse nature of user-generated information poses a significant challenge, limiting the accurate extraction of relevant temporal information and the effective fusion of data across multiple modalities. This paper introduces MMFformer, a multimodal depression detection network designed to retrieve depressive spatio-temporal high-level patterns from multimodal social media information. The transformer network with residual connections captures spatial features from videos, and a transformer encoder is exploited to design important temporal dynamics in audio. Moreover, the fusion architecture fused the extracted features through late and intermediate fusion strategies to find out the most relevant intermodal correlations among them. Finally, the proposed network is assessed on two large-scale depression detection datasets, and the results clearly reveal that it surpasses existing state-of-the-art approaches, improving the F1-Score by 13.92% for D-Vlog dataset and 7.74% for LMVD dataset. The code is made available publicly at https://github.com/rezwanh001/Large-Scale-Multimodal-Depression-Detection.
title MMFformer: Multimodal Fusion Transformer Network for Depression Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.06701