Otter: A Multi-Modal Model with In-Context Instruction Tuning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Bo, Zhang, Yuanhan, Chen, Liangyu, Wang, Jinghao, Pu, Fanyi, Cahyono, Joshua Adrian, Yang, Jingkang, Liu, Ziwei
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911078678528000
author Li, Bo
Zhang, Yuanhan
Chen, Liangyu
Wang, Jinghao
Pu, Fanyi
Cahyono, Joshua Adrian
Yang, Jingkang
Liu, Ziwei
author_facet Li, Bo
Zhang, Yuanhan
Chen, Liangyu
Wang, Jinghao
Pu, Fanyi
Cahyono, Joshua Adrian
Yang, Jingkang
Liu, Ziwei
contents Recent advances in Large Multimodal Models (LMMs) have unveiled great potential as visual assistants. However, most existing works focus on responding to individual instructions or using previous dialogues for contextual understanding. There is little discussion on employing both images and text as in-context examples to enhance the instruction following capability. To bridge this gap, we introduce the \textbf{Otter} model to leverage both textual and visual in-context examples for instruction tuning. Specifically, Otter builds upon Flamingo with Perceiver architecture, and has been instruction tuned for general purpose multi-modal assistant. Otter seamlessly processes multi-modal inputs, supporting modalities including text, multiple images, and dynamic video content. To support the training of Otter, we present the \textbf{MIMIC-IT} (\textbf{M}ult\textbf{I}-\textbf{M}odal \textbf{I}n-\textbf{C}ontext \textbf{I}nstruction \textbf{T}uning) dataset, which encompasses over 3 million multi-modal instruction-response pairs, including approximately 2.2 million unique instructions across a broad spectrum of images and videos. MIMIC-IT has been carefully curated to feature a diverse array of in-context examples for each entry. Comprehensive evaluations suggest that instruction tuning with these in-context examples substantially enhances model convergence and generalization capabilities. Notably, the extensive scenario coverage provided by the MIMIC-IT dataset empowers the Otter model to excel in tasks involving complex video and multi-image understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2305_03726
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Otter: A Multi-Modal Model with In-Context Instruction Tuning
Li, Bo
Zhang, Yuanhan
Chen, Liangyu
Wang, Jinghao
Pu, Fanyi
Cahyono, Joshua Adrian
Yang, Jingkang
Liu, Ziwei
Computer Vision and Pattern Recognition
Computation and Language
Recent advances in Large Multimodal Models (LMMs) have unveiled great potential as visual assistants. However, most existing works focus on responding to individual instructions or using previous dialogues for contextual understanding. There is little discussion on employing both images and text as in-context examples to enhance the instruction following capability. To bridge this gap, we introduce the \textbf{Otter} model to leverage both textual and visual in-context examples for instruction tuning. Specifically, Otter builds upon Flamingo with Perceiver architecture, and has been instruction tuned for general purpose multi-modal assistant. Otter seamlessly processes multi-modal inputs, supporting modalities including text, multiple images, and dynamic video content. To support the training of Otter, we present the \textbf{MIMIC-IT} (\textbf{M}ult\textbf{I}-\textbf{M}odal \textbf{I}n-\textbf{C}ontext \textbf{I}nstruction \textbf{T}uning) dataset, which encompasses over 3 million multi-modal instruction-response pairs, including approximately 2.2 million unique instructions across a broad spectrum of images and videos. MIMIC-IT has been carefully curated to feature a diverse array of in-context examples for each entry. Comprehensive evaluations suggest that instruction tuning with these in-context examples substantially enhances model convergence and generalization capabilities. Notably, the extensive scenario coverage provided by the MIMIC-IT dataset empowers the Otter model to excel in tasks involving complex video and multi-image understanding.
title Otter: A Multi-Modal Model with In-Context Instruction Tuning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2305.03726