FLAM: Frame-Wise Language-Audio Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yusong, Tsirigotis, Christos, Chen, Ke, Huang, Cheng-Zhi Anna, Courville, Aaron, Nieto, Oriol, Seetharaman, Prem, Salamon, Justin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910994361483264
author Wu, Yusong
Tsirigotis, Christos
Chen, Ke
Huang, Cheng-Zhi Anna
Courville, Aaron
Nieto, Oriol
Seetharaman, Prem
Salamon, Justin
author_facet Wu, Yusong
Tsirigotis, Christos
Chen, Ke
Huang, Cheng-Zhi Anna
Courville, Aaron
Nieto, Oriol
Seetharaman, Prem
Salamon, Justin
contents Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05335
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FLAM: Frame-Wise Language-Audio Modeling
Wu, Yusong
Tsirigotis, Christos
Chen, Ke
Huang, Cheng-Zhi Anna
Courville, Aaron
Nieto, Oriol
Seetharaman, Prem
Salamon, Justin
Sound
Audio and Speech Processing
Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks.
title FLAM: Frame-Wise Language-Audio Modeling
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.05335