TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913374617468928 |
|---|---|
| author | Li, Huayang Li, Siheng Cai, Deng Wang, Longyue Liu, Lemao Watanabe, Taro Yang, Yujiu Shi, Shuming |
| author_facet | Li, Huayang Li, Siheng Cai, Deng Wang, Longyue Liu, Lemao Watanabe, Taro Yang, Yujiu Shi, Shuming |
| contents | Large language models with instruction-following abilities have revolutionized the field of artificial intelligence. These models show exceptional generalizability to tackle various real-world tasks through their natural language interfaces. However, their performance heavily relies on high-quality exemplar data, which is often difficult to obtain. This challenge is further exacerbated when it comes to multimodal instruction following. We introduce TextBind, an almost annotation-free framework for empowering larger language models with the multi-turn interleaved multimodal instruction-following capabilities. Our approach requires only image-caption pairs and generates multi-turn multimodal instruction-response conversations from a language model. To accommodate interleaved image-text inputs and outputs, we devise MIM, a language model-centric architecture that seamlessly integrates image encoder and decoder models. We release our dataset, model, and demo to foster future research in the area of multimodal instruction following. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_08637 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild Li, Huayang Li, Siheng Cai, Deng Wang, Longyue Liu, Lemao Watanabe, Taro Yang, Yujiu Shi, Shuming Computation and Language Artificial Intelligence Large language models with instruction-following abilities have revolutionized the field of artificial intelligence. These models show exceptional generalizability to tackle various real-world tasks through their natural language interfaces. However, their performance heavily relies on high-quality exemplar data, which is often difficult to obtain. This challenge is further exacerbated when it comes to multimodal instruction following. We introduce TextBind, an almost annotation-free framework for empowering larger language models with the multi-turn interleaved multimodal instruction-following capabilities. Our approach requires only image-caption pairs and generates multi-turn multimodal instruction-response conversations from a language model. To accommodate interleaved image-text inputs and outputs, we devise MIM, a language model-centric architecture that seamlessly integrates image encoder and decoder models. We release our dataset, model, and demo to foster future research in the area of multimodal instruction following. |
| title | TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2309.08637 |