TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910215324041216 |
|---|---|
| author | Kim, Seongah Tran, Dinh Phu Hwang, Hyeontaek Wazir, Saad Minh, Duc Do Kim, Daeyoung |
| author_facet | Kim, Seongah Tran, Dinh Phu Hwang, Hyeontaek Wazir, Saad Minh, Duc Do Kim, Daeyoung |
| contents | Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to use text as a semantic anchor for audio-visual representation learning. To this end, we introduce a parameter-efficient adaptation framework built on frozen audio and visual encoders, centered on Text-Bridged Audio-Visual Adapter (TB-AVA), which enables text-mediated interaction between audio and visual streams. At the core of TB-AVA, Gated Semantic Modulation (GSM) selectively modulates feature channels based on text-inferred semantic relevance. We evaluate the proposed approach on multiple benchmarks, including AVE, AVS, and AVVP, where the proposed framework achieves state-of-the-art performance, demonstrating text as an effective semantic anchor for parameter-efficient fine-tuning (PEFT) in audio-visual learning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11572 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning Kim, Seongah Tran, Dinh Phu Hwang, Hyeontaek Wazir, Saad Minh, Duc Do Kim, Daeyoung Computer Vision and Pattern Recognition Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to use text as a semantic anchor for audio-visual representation learning. To this end, we introduce a parameter-efficient adaptation framework built on frozen audio and visual encoders, centered on Text-Bridged Audio-Visual Adapter (TB-AVA), which enables text-mediated interaction between audio and visual streams. At the core of TB-AVA, Gated Semantic Modulation (GSM) selectively modulates feature channels based on text-inferred semantic relevance. We evaluate the proposed approach on multiple benchmarks, including AVE, AVS, and AVVP, where the proposed framework achieves state-of-the-art performance, demonstrating text as an effective semantic anchor for parameter-efficient fine-tuning (PEFT) in audio-visual learning. |
| title | TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.11572 |