RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Haoqin, Tian, Jingguang, Zhou, Jiaming, Wang, Hui, He, Jiabei, Zhao, Shiwan, Kong, Xiangyu, Hu, Desheng, Xu, Xinkang, Hu, Xinhui, Qin, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909623025401856
author Sun, Haoqin
Tian, Jingguang
Zhou, Jiaming
Wang, Hui
He, Jiabei
Zhao, Shiwan
Kong, Xiangyu
Hu, Desheng
Xu, Xinkang
Hu, Xinhui
Qin, Yong
author_facet Sun, Haoqin
Tian, Jingguang
Zhou, Jiaming
Wang, Hui
He, Jiabei
Zhao, Shiwan
Kong, Xiangyu
Hu, Desheng
Xu, Xinkang
Hu, Xinhui
Qin, Yong
contents The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the emerging field of emotional speaking style description (ESSD), cross-modal contrastive pretraining remains largely unexplored. In this paper, we propose a novel speech retrieval task called emotional speaking style retrieval (ESSR), and ESS-CLAP, an emotional speaking style CLAP model tailored for learning relationship between speech and natural language descriptions. In addition, we further propose relation-augmented CLAP (RA-CLAP) to address the limitation of traditional methods that assume a strict binary relationship between caption and audio. The model leverages self-distillation to learn the potential local matching relationships between speech and descriptions, thereby enhancing generalization ability. The experimental results validate the effectiveness of RA-CLAP, providing valuable reference in ESSD.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19437
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Sun, Haoqin
Tian, Jingguang
Zhou, Jiaming
Wang, Hui
He, Jiabei
Zhao, Shiwan
Kong, Xiangyu
Hu, Desheng
Xu, Xinkang
Hu, Xinhui
Qin, Yong
Sound
Audio and Speech Processing
The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the emerging field of emotional speaking style description (ESSD), cross-modal contrastive pretraining remains largely unexplored. In this paper, we propose a novel speech retrieval task called emotional speaking style retrieval (ESSR), and ESS-CLAP, an emotional speaking style CLAP model tailored for learning relationship between speech and natural language descriptions. In addition, we further propose relation-augmented CLAP (RA-CLAP) to address the limitation of traditional methods that assume a strict binary relationship between caption and audio. The model leverages self-distillation to learn the potential local matching relationships between speech and descriptions, thereby enhancing generalization ability. The experimental results validate the effectiveness of RA-CLAP, providing valuable reference in ESSD.
title RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.19437