Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909322114498560 |
|---|---|
| author | Phukan, Orchid Chetia Behera, Swarup Ranjan Singh, Shubham Singh, Muskaan Rajan, Vandana Buduru, Arun Balaji Sharma, Rajesh Prasanna, S. R. Mahadeva |
| author_facet | Phukan, Orchid Chetia Behera, Swarup Ranjan Singh, Shubham Singh, Muskaan Rajan, Vandana Buduru, Arun Balaji Sharma, Rajesh Prasanna, S. R. Mahadeva |
| contents | In this study, we address the challenge of depression detection from speech, focusing on the potential of non-semantic features (NSFs) to capture subtle markers of depression. While prior research has leveraged various features for this task, NSFs-extracted from pre-trained models (PTMs) designed for non-semantic tasks such as paralinguistic speech processing (TRILLsson), speaker recognition (x-vector), and emotion recognition (emoHuBERT)-have shown significant promise. However, the potential of combining these diverse features has not been fully explored. In this work, we demonstrate that the amalgamation of NSFs results in complementary behavior, leading to enhanced depression detection performance. Furthermore, to our end, we introduce a simple novel framework, FuSeR, designed to effectively combine these features. Our results show that FuSeR outperforms models utilizing individual NSFs as well as baseline fusion techniques and obtains state-of-the-art (SOTA) performance in E-DAIC benchmark with RMSE of 5.51 and MAE of 4.48, establishing it as a robust approach for depression detection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_14312 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection Phukan, Orchid Chetia Behera, Swarup Ranjan Singh, Shubham Singh, Muskaan Rajan, Vandana Buduru, Arun Balaji Sharma, Rajesh Prasanna, S. R. Mahadeva Audio and Speech Processing Sound 68T45 I.2.7 In this study, we address the challenge of depression detection from speech, focusing on the potential of non-semantic features (NSFs) to capture subtle markers of depression. While prior research has leveraged various features for this task, NSFs-extracted from pre-trained models (PTMs) designed for non-semantic tasks such as paralinguistic speech processing (TRILLsson), speaker recognition (x-vector), and emotion recognition (emoHuBERT)-have shown significant promise. However, the potential of combining these diverse features has not been fully explored. In this work, we demonstrate that the amalgamation of NSFs results in complementary behavior, leading to enhanced depression detection performance. Furthermore, to our end, we introduce a simple novel framework, FuSeR, designed to effectively combine these features. Our results show that FuSeR outperforms models utilizing individual NSFs as well as baseline fusion techniques and obtains state-of-the-art (SOTA) performance in E-DAIC benchmark with RMSE of 5.51 and MAE of 4.48, establishing it as a robust approach for depression detection. |
| title | Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection |
| topic | Audio and Speech Processing Sound 68T45 I.2.7 |
| url | https://arxiv.org/abs/2409.14312 |