Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cuong, Dinh Viet, Le, Hoang-Bao, Nguyen, An Pham Ngoc, Zhou, Liting, Gurrin, Cathal
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912428651970560
author Cuong, Dinh Viet
Le, Hoang-Bao
Nguyen, An Pham Ngoc
Zhou, Liting
Gurrin, Cathal
author_facet Cuong, Dinh Viet
Le, Hoang-Bao
Nguyen, An Pham Ngoc
Zhou, Liting
Gurrin, Cathal
contents This paper addresses two main objectives. Firstly, we demonstrate the impressive performance of the LLaVA-NeXT-interleave on 22 datasets across three different tasks: Multi-Image Reasoning, Documents and Knowledge-Based Understanding and Interactive Multi-Modal Communication. Secondly, we add the Dense Channel Integration (DCI) connector to the LLaVA-NeXT-Interleave and compare its performance against the standard model. We find that the standard model achieves the highest overall accuracy, excelling in vision-heavy tasks like VISION, NLVR2, and Fashion200K. Meanwhile, the DCI-enhanced version shows particular strength on datasets requiring deeper semantic coherence or structured change understanding such as MIT-States_PropertyCoherence and SlideVQA. Our results highlight the potential of combining powerful foundation models with plug-and-play techniques for Interleave tasks. The code is available at https://github.com/dinhvietcuong1996/icme25-inova.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
Cuong, Dinh Viet
Le, Hoang-Bao
Nguyen, An Pham Ngoc
Zhou, Liting
Gurrin, Cathal
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
This paper addresses two main objectives. Firstly, we demonstrate the impressive performance of the LLaVA-NeXT-interleave on 22 datasets across three different tasks: Multi-Image Reasoning, Documents and Knowledge-Based Understanding and Interactive Multi-Modal Communication. Secondly, we add the Dense Channel Integration (DCI) connector to the LLaVA-NeXT-Interleave and compare its performance against the standard model. We find that the standard model achieves the highest overall accuracy, excelling in vision-heavy tasks like VISION, NLVR2, and Fashion200K. Meanwhile, the DCI-enhanced version shows particular strength on datasets requiring deeper semantic coherence or structured change understanding such as MIT-States_PropertyCoherence and SlideVQA. Our results highlight the potential of combining powerful foundation models with plug-and-play techniques for Interleave tasks. The code is available at https://github.com/dinhvietcuong1996/icme25-inova.
title Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2506.11737