RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909623025401856 |
|---|---|
| author | Sun, Haoqin Tian, Jingguang Zhou, Jiaming Wang, Hui He, Jiabei Zhao, Shiwan Kong, Xiangyu Hu, Desheng Xu, Xinkang Hu, Xinhui Qin, Yong |
| author_facet | Sun, Haoqin Tian, Jingguang Zhou, Jiaming Wang, Hui He, Jiabei Zhao, Shiwan Kong, Xiangyu Hu, Desheng Xu, Xinkang Hu, Xinhui Qin, Yong |
| contents | The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the emerging field of emotional speaking style description (ESSD), cross-modal contrastive pretraining remains largely unexplored. In this paper, we propose a novel speech retrieval task called emotional speaking style retrieval (ESSR), and ESS-CLAP, an emotional speaking style CLAP model tailored for learning relationship between speech and natural language descriptions. In addition, we further propose relation-augmented CLAP (RA-CLAP) to address the limitation of traditional methods that assume a strict binary relationship between caption and audio. The model leverages self-distillation to learn the potential local matching relationships between speech and descriptions, thereby enhancing generalization ability. The experimental results validate the effectiveness of RA-CLAP, providing valuable reference in ESSD. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_19437 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Sun, Haoqin Tian, Jingguang Zhou, Jiaming Wang, Hui He, Jiabei Zhao, Shiwan Kong, Xiangyu Hu, Desheng Xu, Xinkang Hu, Xinhui Qin, Yong Sound Audio and Speech Processing The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the emerging field of emotional speaking style description (ESSD), cross-modal contrastive pretraining remains largely unexplored. In this paper, we propose a novel speech retrieval task called emotional speaking style retrieval (ESSR), and ESS-CLAP, an emotional speaking style CLAP model tailored for learning relationship between speech and natural language descriptions. In addition, we further propose relation-augmented CLAP (RA-CLAP) to address the limitation of traditional methods that assume a strict binary relationship between caption and audio. The model leverages self-distillation to learn the potential local matching relationships between speech and descriptions, thereby enhancing generalization ability. The experimental results validate the effectiveness of RA-CLAP, providing valuable reference in ESSD. |
| title | RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.19437 |