LLMBind: A Unified Modality-Task Integration Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917227759927296 |
|---|---|
| author | Zhu, Bin Ning, Munan Jin, Peng Lin, Bin Huang, Jinfa Song, Qi Zhang, Junwu Tang, Zhenyu Pan, Mingjun Yuan, Li |
| author_facet | Zhu, Bin Ning, Munan Jin, Peng Lin, Bin Huang, Jinfa Song, Qi Zhang, Junwu Tang, Zhenyu Pan, Mingjun Yuan, Li |
| contents | Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted task extensibility or severe performance degradation due to modality interference. n this paper, we present LLMBind, an extensible framework that unifies multimodal tasks through a dual-pathway mechanism: In-Situ semantic embeddings for localization-sensitive tasks like semantic segmentation and Ex-Situ task-prompts for generation across image, video, and audio modalities. Additionally, we employ a Mixture-of-Experts (MoE) architecture to route task-specific tokens, thereby achieving modality disentanglement and mitigating negative transfer. We also curate a 400k multi-turn interactive dataset focused on iterative visual refinement to enable human-like interaction. Extensive experiments demonstrate that LLMBind achieves excellent performance across multiple perception and generation benchmarks while maintaining superior expandability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_14891 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LLMBind: A Unified Modality-Task Integration Framework Zhu, Bin Ning, Munan Jin, Peng Lin, Bin Huang, Jinfa Song, Qi Zhang, Junwu Tang, Zhenyu Pan, Mingjun Yuan, Li Computation and Language Artificial Intelligence Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted task extensibility or severe performance degradation due to modality interference. n this paper, we present LLMBind, an extensible framework that unifies multimodal tasks through a dual-pathway mechanism: In-Situ semantic embeddings for localization-sensitive tasks like semantic segmentation and Ex-Situ task-prompts for generation across image, video, and audio modalities. Additionally, we employ a Mixture-of-Experts (MoE) architecture to route task-specific tokens, thereby achieving modality disentanglement and mitigating negative transfer. We also curate a 400k multi-turn interactive dataset focused on iterative visual refinement to enable human-like interaction. Extensive experiments demonstrate that LLMBind achieves excellent performance across multiple perception and generation benchmarks while maintaining superior expandability. |
| title | LLMBind: A Unified Modality-Task Integration Framework |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2402.14891 |