Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nagase, Ryotaro, Sumiyoshi, Takashi, Yamashita, Natsuo, Dohi, Kota, Kawaguchi, Yohei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910648146853888
author Nagase, Ryotaro
Sumiyoshi, Takashi
Yamashita, Natsuo
Dohi, Kota
Kawaguchi, Yohei
author_facet Nagase, Ryotaro
Sumiyoshi, Takashi
Yamashita, Natsuo
Dohi, Kota
Kawaguchi, Yohei
contents This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited to recognizing emotions defined by a single word. Moreover, we have the motivation to recognize unknown bipolar emotions such as ``I want to buy - I do not want to buy.'' In order to allow the model to define classes using sentences freely and to estimate unknown bipolar emotions, our proposed method expands upon the contrastive language-audio pre-training (CLAP) framework by introducing multi-class and multi-task settings. We also focus on purchase intention as a bipolar emotion and investigate the model's performance to zero-shot estimate it. This study is the first attempt to estimate purchase intention from speech directly. Experiments confirm that the results of zero-shot estimation by the proposed method are at the same level as those of the model trained by supervised learning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09636
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
Nagase, Ryotaro
Sumiyoshi, Takashi
Yamashita, Natsuo
Dohi, Kota
Kawaguchi, Yohei
Audio and Speech Processing
Artificial Intelligence
Machine Learning
This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited to recognizing emotions defined by a single word. Moreover, we have the motivation to recognize unknown bipolar emotions such as ``I want to buy - I do not want to buy.'' In order to allow the model to define classes using sentences freely and to estimate unknown bipolar emotions, our proposed method expands upon the contrastive language-audio pre-training (CLAP) framework by introducing multi-class and multi-task settings. We also focus on purchase intention as a bipolar emotion and investigate the model's performance to zero-shot estimate it. This study is the first attempt to estimate purchase intention from speech directly. Experiments confirm that the results of zero-shot estimation by the proposed method are at the same level as those of the model trained by supervised learning.
title Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.09636