CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Suharitdamrong, Wish, Alex, Tony, Awais, Muhammad, Ahmed, Sara
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917383668498432
author Suharitdamrong, Wish
Alex, Tony
Awais, Muhammad
Ahmed, Sara
author_facet Suharitdamrong, Wish
Alex, Tony
Awais, Muhammad
Ahmed, Sara
contents Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challenge. Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) enable lightweight adaptation, yet they operate in isolation within each modality, limiting their ability in capturing cross-modal interactions. In this paper, we take a step in bridging this gap with Cross-Modal Low-Rank Adaptation (CoLA), a novel PEFT framework that extends LoRA by introducing a dedicated inter-modal adaptation pathway alongside the standard intra-modal one. This dual-path design enables CoLA to adapt unimodal foundation models to multimodal tasks effectively, without interference between modality-specific and cross-modal learning. We evaluate CoLA across a range of vision-language (RefCOCO, RefCOCO+, RefCOCOg) and audio-visual (AVE, AVS) benchmarks, where it consistently outperforms LORA, achieving a relative gain of around 3\% and 2\%, respectively, while maintaining parameter efficiency. Notably, CoLA enables the first multi-task PEFT framework for visual grounding, bridging a key gap in efficient multimodal adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03314
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
Suharitdamrong, Wish
Alex, Tony
Awais, Muhammad
Ahmed, Sara
Computer Vision and Pattern Recognition
Computation and Language
Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challenge. Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) enable lightweight adaptation, yet they operate in isolation within each modality, limiting their ability in capturing cross-modal interactions. In this paper, we take a step in bridging this gap with Cross-Modal Low-Rank Adaptation (CoLA), a novel PEFT framework that extends LoRA by introducing a dedicated inter-modal adaptation pathway alongside the standard intra-modal one. This dual-path design enables CoLA to adapt unimodal foundation models to multimodal tasks effectively, without interference between modality-specific and cross-modal learning. We evaluate CoLA across a range of vision-language (RefCOCO, RefCOCO+, RefCOCOg) and audio-visual (AVE, AVS) benchmarks, where it consistently outperforms LORA, achieving a relative gain of around 3\% and 2\%, respectively, while maintaining parameter efficiency. Notably, CoLA enables the first multi-task PEFT framework for visual grounding, bridging a key gap in efficient multimodal adaptation.
title CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.03314