FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zehan, Zhang, Ziang, Cheng, Xize, Huang, Rongjie, Liu, Luping, Ye, Zhenhui, Huang, Haifeng, Zhao, Yang, Jin, Tao, Gao, Peng, Zhao, Zhou
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911872483065856
author Wang, Zehan
Zhang, Ziang
Cheng, Xize
Huang, Rongjie
Liu, Luping
Ye, Zhenhui
Huang, Haifeng
Zhao, Yang
Jin, Tao
Gao, Peng
Zhao, Zhou
author_facet Wang, Zehan
Zhang, Ziang
Cheng, Xize
Huang, Rongjie
Liu, Luping
Ye, Zhenhui
Huang, Haifeng
Zhao, Yang
Jin, Tao
Gao, Peng
Zhao, Zhou
contents Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that treats multimodal representation spaces as basic units, and freely augments pre-trained unified space by integrating knowledge from extra expert spaces via "space bonds". Specifically, we introduce two kinds of basic space bonds: 1) Space Displacement Bond and 2) Space Combination Bond. Based on these basic bonds, we design Complex Sequential & Parallel Bonds to effectively integrate multiple spaces simultaneously. Benefiting from the modularization concept, we further propose a coarse-to-fine customized inference strategy to flexibly adjust the enhanced unified space for different purposes. Experimentally, we bind ImageBind with extra image-text and audio-text expert spaces, resulting in three main variants: ImageBind++, InternVL_IB, and InternVL_IB++. These resulting spaces outperform ImageBind on 5 audio-image-text downstream tasks across 9 datasets. Moreover, via customized inference, it even surpasses the advanced audio-text and image-text expert spaces.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04883
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
Wang, Zehan
Zhang, Ziang
Cheng, Xize
Huang, Rongjie
Liu, Luping
Ye, Zhenhui
Huang, Haifeng
Zhao, Yang
Jin, Tao
Gao, Peng
Zhao, Zhou
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that treats multimodal representation spaces as basic units, and freely augments pre-trained unified space by integrating knowledge from extra expert spaces via "space bonds". Specifically, we introduce two kinds of basic space bonds: 1) Space Displacement Bond and 2) Space Combination Bond. Based on these basic bonds, we design Complex Sequential & Parallel Bonds to effectively integrate multiple spaces simultaneously. Benefiting from the modularization concept, we further propose a coarse-to-fine customized inference strategy to flexibly adjust the enhanced unified space for different purposes. Experimentally, we bind ImageBind with extra image-text and audio-text expert spaces, resulting in three main variants: ImageBind++, InternVL_IB, and InternVL_IB++. These resulting spaces outperform ImageBind on 5 audio-image-text downstream tasks across 9 datasets. Moreover, via customized inference, it even surpasses the advanced audio-text and image-text expert spaces.
title FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.04883