OneLLM: One Framework to Align All Modalities with Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Jiaming, Gong, Kaixiong, Zhang, Yiyuan, Wang, Jiaqi, Zhang, Kaipeng, Lin, Dahua, Qiao, Yu, Gao, Peng, Yue, Xiangyu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913641771565056
author Han, Jiaming
Gong, Kaixiong
Zhang, Yiyuan
Wang, Jiaqi
Zhang, Kaipeng
Lin, Dahua
Qiao, Yu
Gao, Peng
Yue, Xiangyu
author_facet Han, Jiaming
Gong, Kaixiong
Zhang, Yiyuan
Wang, Jiaqi
Zhang, Kaipeng
Lin, Dahua
Qiao, Yu
Gao, Peng
Yue, Xiangyu
contents Multimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However, existing works rely heavily on modality-specific encoders, which usually differ in architecture and are limited to common modalities. In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using a unified framework. We achieve this through a unified multimodal encoder and a progressive multimodal alignment pipeline. In detail, we first train an image projection module to connect a vision encoder with LLM. Then, we build a universal projection module (UPM) by mixing multiple image projection modules and dynamic routing. Finally, we progressively align more modalities to LLM with the UPM. To fully leverage the potential of OneLLM in following instructions, we also curated a comprehensive multimodal instruction dataset, including 2M items from image, audio, video, point cloud, depth/normal map, IMU and fMRI brain activity. OneLLM is evaluated on 25 diverse benchmarks, encompassing tasks such as multimodal captioning, question answering and reasoning, where it delivers excellent performance. Code, data, model and online demo are available at https://github.com/csuhan/OneLLM
format Preprint
id arxiv_https___arxiv_org_abs_2312_03700
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle OneLLM: One Framework to Align All Modalities with Language
Han, Jiaming
Gong, Kaixiong
Zhang, Yiyuan
Wang, Jiaqi
Zhang, Kaipeng
Lin, Dahua
Qiao, Yu
Gao, Peng
Yue, Xiangyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
Multimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However, existing works rely heavily on modality-specific encoders, which usually differ in architecture and are limited to common modalities. In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using a unified framework. We achieve this through a unified multimodal encoder and a progressive multimodal alignment pipeline. In detail, we first train an image projection module to connect a vision encoder with LLM. Then, we build a universal projection module (UPM) by mixing multiple image projection modules and dynamic routing. Finally, we progressively align more modalities to LLM with the UPM. To fully leverage the potential of OneLLM in following instructions, we also curated a comprehensive multimodal instruction dataset, including 2M items from image, audio, video, point cloud, depth/normal map, IMU and fMRI brain activity. OneLLM is evaluated on 25 diverse benchmarks, encompassing tasks such as multimodal captioning, question answering and reasoning, where it delivers excellent performance. Code, data, model and online demo are available at https://github.com/csuhan/OneLLM
title OneLLM: One Framework to Align All Modalities with Language
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
url https://arxiv.org/abs/2312.03700