TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Seongah, Tran, Dinh Phu, Hwang, Hyeontaek, Wazir, Saad, Minh, Duc Do, Kim, Daeyoung
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910215324041216
author Kim, Seongah
Tran, Dinh Phu
Hwang, Hyeontaek
Wazir, Saad
Minh, Duc Do
Kim, Daeyoung
author_facet Kim, Seongah
Tran, Dinh Phu
Hwang, Hyeontaek
Wazir, Saad
Minh, Duc Do
Kim, Daeyoung
contents Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to use text as a semantic anchor for audio-visual representation learning. To this end, we introduce a parameter-efficient adaptation framework built on frozen audio and visual encoders, centered on Text-Bridged Audio-Visual Adapter (TB-AVA), which enables text-mediated interaction between audio and visual streams. At the core of TB-AVA, Gated Semantic Modulation (GSM) selectively modulates feature channels based on text-inferred semantic relevance. We evaluate the proposed approach on multiple benchmarks, including AVE, AVS, and AVVP, where the proposed framework achieves state-of-the-art performance, demonstrating text as an effective semantic anchor for parameter-efficient fine-tuning (PEFT) in audio-visual learning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11572
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
Kim, Seongah
Tran, Dinh Phu
Hwang, Hyeontaek
Wazir, Saad
Minh, Duc Do
Kim, Daeyoung
Computer Vision and Pattern Recognition
Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to use text as a semantic anchor for audio-visual representation learning. To this end, we introduce a parameter-efficient adaptation framework built on frozen audio and visual encoders, centered on Text-Bridged Audio-Visual Adapter (TB-AVA), which enables text-mediated interaction between audio and visual streams. At the core of TB-AVA, Gated Semantic Modulation (GSM) selectively modulates feature channels based on text-inferred semantic relevance. We evaluate the proposed approach on multiple benchmarks, including AVE, AVS, and AVVP, where the proposed framework achieves state-of-the-art performance, demonstrating text as an effective semantic anchor for parameter-efficient fine-tuning (PEFT) in audio-visual learning.
title TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.11572