Human-CLAP: Human-perception-based contrastive language-audio pretraining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Takano, Taisei, Okamoto, Yuki, Kanamori, Yusuke, Saito, Yuki, Nagase, Ryotaro, Saruwatari, Hiroshi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917326502232064
author Takano, Taisei
Okamoto, Yuki
Kanamori, Yusuke
Saito, Yuki
Nagase, Ryotaro
Saruwatari, Hiroshi
author_facet Takano, Taisei
Okamoto, Yuki
Kanamori, Yusuke
Saito, Yuki
Nagase, Ryotaro
Saruwatari, Hiroshi
contents Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23553
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Human-CLAP: Human-perception-based contrastive language-audio pretraining
Takano, Taisei
Okamoto, Yuki
Kanamori, Yusuke
Saito, Yuki
Nagase, Ryotaro
Saruwatari, Hiroshi
Audio and Speech Processing
Sound
Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP.
title Human-CLAP: Human-perception-based contrastive language-audio pretraining
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.23553