Aligning Audio Captions with Human Preferences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hegde, Kartik, Mahfuz, Rehana, Guo, Yinyi, Visser, Erik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911466614947840
author Hegde, Kartik
Mahfuz, Rehana
Guo, Yinyi
Visser, Erik
author_facet Hegde, Kartik
Mahfuz, Rehana
Guo, Yinyi
Visser, Erik
contents Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios. To address this, we propose a preference-aligned audio captioning framework based on Reinforcement Learning from Human Feedback (RLHF). To capture nuanced preferences, we train a Contrastive Language-Audio Pretraining (CLAP) based reward model using human-labeled pairwise preference data. This reward model is integrated into an RL framework to fine-tune any baseline captioning system without ground-truth annotations. Extensive human evaluations across multiple datasets show that our method produces captions preferred over baseline models, particularly when baselines fail to provide correct and natural captions. Furthermore, our framework achieves performance comparable to supervised approaches with ground-truth data, demonstrating effective alignment with human preferences and scalability in real-world use.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14659
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning Audio Captions with Human Preferences
Hegde, Kartik
Mahfuz, Rehana
Guo, Yinyi
Visser, Erik
Audio and Speech Processing
Machine Learning
Sound
Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios. To address this, we propose a preference-aligned audio captioning framework based on Reinforcement Learning from Human Feedback (RLHF). To capture nuanced preferences, we train a Contrastive Language-Audio Pretraining (CLAP) based reward model using human-labeled pairwise preference data. This reward model is integrated into an RL framework to fine-tune any baseline captioning system without ground-truth annotations. Extensive human evaluations across multiple datasets show that our method produces captions preferred over baseline models, particularly when baselines fail to provide correct and natural captions. Furthermore, our framework achieves performance comparable to supervised approaches with ground-truth data, demonstrating effective alignment with human preferences and scalability in real-world use.
title Aligning Audio Captions with Human Preferences
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2509.14659