HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Runhui, Ding, Xinpeng, Wang, Chunwei, Han, Jianhua, Liu, Yulong, Zhao, Hengshuang, Xu, Hang, Hou, Lu, Zhang, Wei, Liang, Xiaodan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908755889750016
author Huang, Runhui
Ding, Xinpeng
Wang, Chunwei
Han, Jianhua
Liu, Yulong
Zhao, Hengshuang
Xu, Hang
Hou, Lu
Zhang, Wei
Liang, Xiaodan
author_facet Huang, Runhui
Ding, Xinpeng
Wang, Chunwei
Han, Jianhua
Liu, Yulong
Zhao, Hengshuang
Xu, Hang
Hou, Lu
Zhang, Wei
Liang, Xiaodan
contents High-resolution inputs enable Large Vision-Language Models (LVLMs) to discern finer visual details, enhancing their comprehension capabilities. To reduce the training and computation costs caused by high-resolution input, one promising direction is to use sliding windows to slice the input into uniform patches, each matching the input size of the well-trained vision encoder. Although efficient, this slicing strategy leads to the fragmentation of original input, i.e., the continuity of contextual information and spatial geometry is lost across patches, adversely affecting performance in cross-patch context perception and position-specific tasks. To overcome these shortcomings, we introduce HiRes-LLaVA, a novel framework designed to efficiently process any size of high-resolution input without altering the original contextual and geometric information. HiRes-LLaVA comprises two innovative components: (i) a SliceRestore adapter that reconstructs sliced patches into their original form, efficiently extracting both global and local features via down-up-sampling and convolution layers, and (ii) a Self-Mining Sampler to compresses the vision tokens based on themselves, preserving the original context and positional information while reducing training overhead. To assess the ability of handling context fragmentation, we construct a new benchmark, EntityGrid-QA, consisting of edge-related and position-related tasks. Our comprehensive experiments demonstrate the superiority of HiRes-LLaVA on both existing public benchmarks and on EntityGrid-QA, particularly on document-oriented tasks, establishing new standards for handling high-resolution inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08706
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
Huang, Runhui
Ding, Xinpeng
Wang, Chunwei
Han, Jianhua
Liu, Yulong
Zhao, Hengshuang
Xu, Hang
Hou, Lu
Zhang, Wei
Liang, Xiaodan
Computer Vision and Pattern Recognition
High-resolution inputs enable Large Vision-Language Models (LVLMs) to discern finer visual details, enhancing their comprehension capabilities. To reduce the training and computation costs caused by high-resolution input, one promising direction is to use sliding windows to slice the input into uniform patches, each matching the input size of the well-trained vision encoder. Although efficient, this slicing strategy leads to the fragmentation of original input, i.e., the continuity of contextual information and spatial geometry is lost across patches, adversely affecting performance in cross-patch context perception and position-specific tasks. To overcome these shortcomings, we introduce HiRes-LLaVA, a novel framework designed to efficiently process any size of high-resolution input without altering the original contextual and geometric information. HiRes-LLaVA comprises two innovative components: (i) a SliceRestore adapter that reconstructs sliced patches into their original form, efficiently extracting both global and local features via down-up-sampling and convolution layers, and (ii) a Self-Mining Sampler to compresses the vision tokens based on themselves, preserving the original context and positional information while reducing training overhead. To assess the ability of handling context fragmentation, we construct a new benchmark, EntityGrid-QA, consisting of edge-related and position-related tasks. Our comprehensive experiments demonstrate the superiority of HiRes-LLaVA on both existing public benchmarks and on EntityGrid-QA, particularly on document-oriented tasks, establishing new standards for handling high-resolution inputs.
title HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.08706