Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Xiao, Niu, Tianhao, Xie, Yuxi, Qin, Libo, Che, Wanxiang, Kan, Min-Yen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929619943292928
author Xu, Xiao
Niu, Tianhao
Xie, Yuxi
Qin, Libo
Che, Wanxiang
Kan, Min-Yen
author_facet Xu, Xiao
Niu, Tianhao
Xie, Yuxi
Qin, Libo
Che, Wanxiang
Kan, Min-Yen
contents Multimodal Large Language Models (MLLMs) excel in vision--language tasks by pre-training solely on coarse-grained concept annotations (e.g., image captions). We hypothesize that integrating fine-grained concept annotations (e.g., object labels and object regions) will further improve performance, as both data granularities complement each other in terms of breadth and depth in concept representation. We introduce a new dataset featuring Multimodal Multi-Grained Concept annotations (MMGiC) for MLLMs. In constructing MMGiC, we explore the impact of different data recipes on multimodal comprehension and generation. Our analyses reveal that multi-grained concept annotations integrate and complement each other, under our structured template and a general MLLM framework. We clearly explore and demonstrate the potential of MMGiC to help MLLMs better locate and learn concepts, aligning vision and language at multiple granularities. We further validate our hypothesis by investigating the fair comparison and effective collaboration between MMGiC and image--caption data on 12 multimodal comprehension and generation benchmarks, e.g., their appropriate combination achieve 3.95% and 2.34% absolute improvements over image--caption data alone on POPE and SEED-Bench. Code, data and models will be available at https://github.com/LooperXX/MMGiC.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
Xu, Xiao
Niu, Tianhao
Xie, Yuxi
Qin, Libo
Che, Wanxiang
Kan, Min-Yen
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Multimodal Large Language Models (MLLMs) excel in vision--language tasks by pre-training solely on coarse-grained concept annotations (e.g., image captions). We hypothesize that integrating fine-grained concept annotations (e.g., object labels and object regions) will further improve performance, as both data granularities complement each other in terms of breadth and depth in concept representation. We introduce a new dataset featuring Multimodal Multi-Grained Concept annotations (MMGiC) for MLLMs. In constructing MMGiC, we explore the impact of different data recipes on multimodal comprehension and generation. Our analyses reveal that multi-grained concept annotations integrate and complement each other, under our structured template and a general MLLM framework. We clearly explore and demonstrate the potential of MMGiC to help MLLMs better locate and learn concepts, aligning vision and language at multiple granularities. We further validate our hypothesis by investigating the fair comparison and effective collaboration between MMGiC and image--caption data on 12 multimodal comprehension and generation benchmarks, e.g., their appropriate combination achieve 3.95% and 2.34% absolute improvements over image--caption data alone on POPE and SEED-Bench. Code, data and models will be available at https://github.com/LooperXX/MMGiC.
title Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.05939