Human-CLAP: Human-perception-based contrastive language-audio pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917326502232064 |
|---|---|
| author | Takano, Taisei Okamoto, Yuki Kanamori, Yusuke Saito, Yuki Nagase, Ryotaro Saruwatari, Hiroshi |
| author_facet | Takano, Taisei Okamoto, Yuki Kanamori, Yusuke Saito, Yuki Nagase, Ryotaro Saruwatari, Hiroshi |
| contents | Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_23553 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Human-CLAP: Human-perception-based contrastive language-audio pretraining Takano, Taisei Okamoto, Yuki Kanamori, Yusuke Saito, Yuki Nagase, Ryotaro Saruwatari, Hiroshi Audio and Speech Processing Sound Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP. |
| title | Human-CLAP: Human-perception-based contrastive language-audio pretraining |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2506.23553 |