Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910648146853888 |
|---|---|
| author | Nagase, Ryotaro Sumiyoshi, Takashi Yamashita, Natsuo Dohi, Kota Kawaguchi, Yohei |
| author_facet | Nagase, Ryotaro Sumiyoshi, Takashi Yamashita, Natsuo Dohi, Kota Kawaguchi, Yohei |
| contents | This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited to recognizing emotions defined by a single word. Moreover, we have the motivation to recognize unknown bipolar emotions such as ``I want to buy - I do not want to buy.'' In order to allow the model to define classes using sentences freely and to estimate unknown bipolar emotions, our proposed method expands upon the contrastive language-audio pre-training (CLAP) framework by introducing multi-class and multi-task settings. We also focus on purchase intention as a bipolar emotion and investigate the model's performance to zero-shot estimate it. This study is the first attempt to estimate purchase intention from speech directly. Experiments confirm that the results of zero-shot estimation by the proposed method are at the same level as those of the model trained by supervised learning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_09636 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition? Nagase, Ryotaro Sumiyoshi, Takashi Yamashita, Natsuo Dohi, Kota Kawaguchi, Yohei Audio and Speech Processing Artificial Intelligence Machine Learning This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited to recognizing emotions defined by a single word. Moreover, we have the motivation to recognize unknown bipolar emotions such as ``I want to buy - I do not want to buy.'' In order to allow the model to define classes using sentences freely and to estimate unknown bipolar emotions, our proposed method expands upon the contrastive language-audio pre-training (CLAP) framework by introducing multi-class and multi-task settings. We also focus on purchase intention as a bipolar emotion and investigate the model's performance to zero-shot estimate it. This study is the first attempt to estimate purchase intention from speech directly. Experiments confirm that the results of zero-shot estimation by the proposed method are at the same level as those of the model trained by supervised learning. |
| title | Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition? |
| topic | Audio and Speech Processing Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2410.09636 |