FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial Landmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ruiqi, Huang, Jinyang, Zhang, Jie, Liu, Xin, Zhang, Xiang, Liu, Zhi, Zhao, Peng, Chen, Sigui, Sun, Xiao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917742857158656
author Wang, Ruiqi
Huang, Jinyang
Zhang, Jie
Liu, Xin
Zhang, Xiang
Liu, Zhi
Zhao, Peng
Chen, Sigui
Sun, Xiao
author_facet Wang, Ruiqi
Huang, Jinyang
Zhang, Jie
Liu, Xin
Zhang, Xiang
Liu, Zhi
Zhao, Peng
Chen, Sigui
Sun, Xiao
contents Depression is a prevalent mental health disorder that significantly impacts individuals' lives and well-being. Early detection and intervention are crucial for effective treatment and management of depression. Recently, there are many end-to-end deep learning methods leveraging the facial expression features for automatic depression detection. However, most current methods overlook the temporal dynamics of facial expressions. Although very recent 3DCNN methods remedy this gap, they introduce more computational cost due to the selection of CNN-based backbones and redundant facial features. To address the above limitations, by considering the timing correlation of facial expressions, we propose a novel framework called FacialPulse, which recognizes depression with high accuracy and speed. By harnessing the bidirectional nature and proficiently addressing long-term dependencies, the Facial Motion Modeling Module (FMMM) is designed in FacialPulse to fully capture temporal features. Since the proposed FMMM has parallel processing capabilities and has the gate mechanism to mitigate gradient vanishing, this module can also significantly boost the training speed. Besides, to effectively use facial landmarks to replace original images to decrease information redundancy, a Facial Landmark Calibration Module (FLCM) is designed to eliminate facial landmark errors to further improve recognition accuracy. Extensive experiments on the AVEC2014 dataset and MMDA dataset (a depression dataset) demonstrate the superiority of FacialPulse on recognition accuracy and speed, with the average MAE (Mean Absolute Error) decreased by 21% compared to baselines, and the recognition speed increased by 100% compared to state-of-the-art methods. Codes are released at https://github.com/volatileee/FacialPulse.
format Preprint
id arxiv_https___arxiv_org_abs_2408_03499
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial Landmarks
Wang, Ruiqi
Huang, Jinyang
Zhang, Jie
Liu, Xin
Zhang, Xiang
Liu, Zhi
Zhao, Peng
Chen, Sigui
Sun, Xiao
Computer Vision and Pattern Recognition
Depression is a prevalent mental health disorder that significantly impacts individuals' lives and well-being. Early detection and intervention are crucial for effective treatment and management of depression. Recently, there are many end-to-end deep learning methods leveraging the facial expression features for automatic depression detection. However, most current methods overlook the temporal dynamics of facial expressions. Although very recent 3DCNN methods remedy this gap, they introduce more computational cost due to the selection of CNN-based backbones and redundant facial features. To address the above limitations, by considering the timing correlation of facial expressions, we propose a novel framework called FacialPulse, which recognizes depression with high accuracy and speed. By harnessing the bidirectional nature and proficiently addressing long-term dependencies, the Facial Motion Modeling Module (FMMM) is designed in FacialPulse to fully capture temporal features. Since the proposed FMMM has parallel processing capabilities and has the gate mechanism to mitigate gradient vanishing, this module can also significantly boost the training speed. Besides, to effectively use facial landmarks to replace original images to decrease information redundancy, a Facial Landmark Calibration Module (FLCM) is designed to eliminate facial landmark errors to further improve recognition accuracy. Extensive experiments on the AVEC2014 dataset and MMDA dataset (a depression dataset) demonstrate the superiority of FacialPulse on recognition accuracy and speed, with the average MAE (Mean Absolute Error) decreased by 21% compared to baselines, and the recognition speed increased by 100% compared to state-of-the-art methods. Codes are released at https://github.com/volatileee/FacialPulse.
title FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial Landmarks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.03499