Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Lincan, Li, Shuang, Ma, Wenxuan, Kang, Jingxuan, Xie, Binhui, Sun, Zixun, Zhu, Chengwei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917692761440256
author Cai, Lincan
Li, Shuang
Ma, Wenxuan
Kang, Jingxuan
Xie, Binhui
Sun, Zixun
Zhu, Chengwei
author_facet Cai, Lincan
Li, Shuang
Ma, Wenxuan
Kang, Jingxuan
Xie, Binhui
Sun, Zixun
Zhu, Chengwei
contents Large-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scarcity of labeled data. In this paper, we propose an end-to-end method, PaRe, to enhance cross-modal fine-tuning, aiming to transfer a large-scale pretrained model to various target modalities. PaRe employs a gating mechanism to select key patches from both source and target data. Through a modality-agnostic Patch Replacement scheme, these patches are preserved and combined to construct data-rich intermediate modalities ranging from easy to hard. By gradually intermediate modality generation, we can not only effectively bridge the modality gap to enhance stability and transferability of cross-modal fine-tuning, but also address the challenge of limited data in the target modality by leveraging enriched intermediate modality data. Compared with hand-designed, general-purpose, task-specific, and state-of-the-art cross-modal fine-tuning approaches, PaRe demonstrates superior performance across three challenging benchmarks, encompassing more than ten modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09003
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
Cai, Lincan
Li, Shuang
Ma, Wenxuan
Kang, Jingxuan
Xie, Binhui
Sun, Zixun
Zhu, Chengwei
Computer Vision and Pattern Recognition
Machine Learning
Large-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scarcity of labeled data. In this paper, we propose an end-to-end method, PaRe, to enhance cross-modal fine-tuning, aiming to transfer a large-scale pretrained model to various target modalities. PaRe employs a gating mechanism to select key patches from both source and target data. Through a modality-agnostic Patch Replacement scheme, these patches are preserved and combined to construct data-rich intermediate modalities ranging from easy to hard. By gradually intermediate modality generation, we can not only effectively bridge the modality gap to enhance stability and transferability of cross-modal fine-tuning, but also address the challenge of limited data in the target modality by leveraging enriched intermediate modality data. Compared with hand-designed, general-purpose, task-specific, and state-of-the-art cross-modal fine-tuning approaches, PaRe demonstrates superior performance across three challenging benchmarks, encompassing more than ten modalities.
title Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.09003