Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pramanick, Shraman, Han, Guangxing, Hou, Rui, Nag, Sayan, Lim, Ser-Nam, Ballas, Nicolas, Wang, Qifan, Chellappa, Rama, Almahairi, Amjad
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913397686140928
author Pramanick, Shraman
Han, Guangxing
Hou, Rui
Nag, Sayan
Lim, Ser-Nam
Ballas, Nicolas
Wang, Qifan
Chellappa, Rama
Almahairi, Amjad
author_facet Pramanick, Shraman
Han, Guangxing
Hou, Rui
Nag, Sayan
Lim, Ser-Nam
Ballas, Nicolas
Wang, Qifan
Chellappa, Rama
Almahairi, Amjad
contents The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems, unifying various vision-language (VL) tasks by instruction tuning. However, due to the enormous diversity in input-output formats in the vision domain, existing general-purpose models fail to successfully integrate segmentation and multi-image inputs with coarse-level tasks into a single framework. In this work, we introduce VistaLLM, a powerful visual system that addresses coarse- and fine-grained VL tasks over single and multiple input images using a unified framework. VistaLLM utilizes an instruction-guided image tokenizer that filters global embeddings using task descriptions to extract compressed and refined features from numerous images. Moreover, VistaLLM employs a gradient-aware adaptive sampling technique to represent binary segmentation masks as sequences, significantly improving over previously used uniform sampling. To bolster the desired capability of VistaLLM, we curate CoinIt, a comprehensive coarse-to-fine instruction tuning dataset with 6.8M samples. We also address the lack of multi-image grounding datasets by introducing a novel task, AttCoSeg (Attribute-level Co-Segmentation), which boosts the model's reasoning and grounding capability over multiple input images. Extensive experiments on a wide range of V- and VL tasks demonstrate the effectiveness of VistaLLM by achieving consistent state-of-the-art performance over strong baselines across all downstream tasks. Our project page can be found at https://shramanpramanick.github.io/VistaLLM/.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12423
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
Pramanick, Shraman
Han, Guangxing
Hou, Rui
Nag, Sayan
Lim, Ser-Nam
Ballas, Nicolas
Wang, Qifan
Chellappa, Rama
Almahairi, Amjad
Computer Vision and Pattern Recognition
Artificial Intelligence
The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems, unifying various vision-language (VL) tasks by instruction tuning. However, due to the enormous diversity in input-output formats in the vision domain, existing general-purpose models fail to successfully integrate segmentation and multi-image inputs with coarse-level tasks into a single framework. In this work, we introduce VistaLLM, a powerful visual system that addresses coarse- and fine-grained VL tasks over single and multiple input images using a unified framework. VistaLLM utilizes an instruction-guided image tokenizer that filters global embeddings using task descriptions to extract compressed and refined features from numerous images. Moreover, VistaLLM employs a gradient-aware adaptive sampling technique to represent binary segmentation masks as sequences, significantly improving over previously used uniform sampling. To bolster the desired capability of VistaLLM, we curate CoinIt, a comprehensive coarse-to-fine instruction tuning dataset with 6.8M samples. We also address the lack of multi-image grounding datasets by introducing a novel task, AttCoSeg (Attribute-level Co-Segmentation), which boosts the model's reasoning and grounding capability over multiple input images. Extensive experiments on a wide range of V- and VL tasks demonstrate the effectiveness of VistaLLM by achieving consistent state-of-the-art performance over strong baselines across all downstream tasks. Our project page can be found at https://shramanpramanick.github.io/VistaLLM/.
title Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2312.12423