LLMBind: A Unified Modality-Task Integration Framework

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Bin, Ning, Munan, Jin, Peng, Lin, Bin, Huang, Jinfa, Song, Qi, Zhang, Junwu, Tang, Zhenyu, Pan, Mingjun, Yuan, Li
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917227759927296
author Zhu, Bin
Ning, Munan
Jin, Peng
Lin, Bin
Huang, Jinfa
Song, Qi
Zhang, Junwu
Tang, Zhenyu
Pan, Mingjun
Yuan, Li
author_facet Zhu, Bin
Ning, Munan
Jin, Peng
Lin, Bin
Huang, Jinfa
Song, Qi
Zhang, Junwu
Tang, Zhenyu
Pan, Mingjun
Yuan, Li
contents Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted task extensibility or severe performance degradation due to modality interference. n this paper, we present LLMBind, an extensible framework that unifies multimodal tasks through a dual-pathway mechanism: In-Situ semantic embeddings for localization-sensitive tasks like semantic segmentation and Ex-Situ task-prompts for generation across image, video, and audio modalities. Additionally, we employ a Mixture-of-Experts (MoE) architecture to route task-specific tokens, thereby achieving modality disentanglement and mitigating negative transfer. We also curate a 400k multi-turn interactive dataset focused on iterative visual refinement to enable human-like interaction. Extensive experiments demonstrate that LLMBind achieves excellent performance across multiple perception and generation benchmarks while maintaining superior expandability.
format Preprint
id arxiv_https___arxiv_org_abs_2402_14891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLMBind: A Unified Modality-Task Integration Framework
Zhu, Bin
Ning, Munan
Jin, Peng
Lin, Bin
Huang, Jinfa
Song, Qi
Zhang, Junwu
Tang, Zhenyu
Pan, Mingjun
Yuan, Li
Computation and Language
Artificial Intelligence
Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted task extensibility or severe performance degradation due to modality interference. n this paper, we present LLMBind, an extensible framework that unifies multimodal tasks through a dual-pathway mechanism: In-Situ semantic embeddings for localization-sensitive tasks like semantic segmentation and Ex-Situ task-prompts for generation across image, video, and audio modalities. Additionally, we employ a Mixture-of-Experts (MoE) architecture to route task-specific tokens, thereby achieving modality disentanglement and mitigating negative transfer. We also curate a 400k multi-turn interactive dataset focused on iterative visual refinement to enable human-like interaction. Extensive experiments demonstrate that LLMBind achieves excellent performance across multiple perception and generation benchmarks while maintaining superior expandability.
title LLMBind: A Unified Modality-Task Integration Framework
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.14891