FLAM: Frame-Wise Language-Audio Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910994361483264 |
|---|---|
| author | Wu, Yusong Tsirigotis, Christos Chen, Ke Huang, Cheng-Zhi Anna Courville, Aaron Nieto, Oriol Seetharaman, Prem Salamon, Justin |
| author_facet | Wu, Yusong Tsirigotis, Christos Chen, Ke Huang, Cheng-Zhi Anna Courville, Aaron Nieto, Oriol Seetharaman, Prem Salamon, Justin |
| contents | Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_05335 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FLAM: Frame-Wise Language-Audio Modeling Wu, Yusong Tsirigotis, Christos Chen, Ke Huang, Cheng-Zhi Anna Courville, Aaron Nieto, Oriol Seetharaman, Prem Salamon, Justin Sound Audio and Speech Processing Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks. |
| title | FLAM: Frame-Wise Language-Audio Modeling |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.05335 |