OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tao, Li, Xiangtai, Fei, Hao, Yuan, Haobo, Wu, Shengqiong, Ji, Shunping, Loy, Chen Change, Yan, Shuicheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916416651788288
author Zhang, Tao
Li, Xiangtai
Fei, Hao
Yuan, Haobo
Wu, Shengqiong
Ji, Shunping
Loy, Chen Change
Yan, Shuicheng
author_facet Zhang, Tao
Li, Xiangtai
Fei, Hao
Yuan, Haobo
Wu, Shengqiong
Ji, Shunping
Loy, Chen Change
Yan, Shuicheng
contents Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user's text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19389
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Zhang, Tao
Li, Xiangtai
Fei, Hao
Yuan, Haobo
Wu, Shengqiong
Ji, Shunping
Loy, Chen Change
Yan, Shuicheng
Computer Vision and Pattern Recognition
Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user's text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research.
title OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.19389