Benchmarking Large and Small MLLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Feng, Xuelu, Li, Yunsheng, Chen, Dongdong, Gao, Mei, Liu, Mengchen, Yuan, Junsong, Qiao, Chunming
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917887062573056
author Feng, Xuelu
Li, Yunsheng
Chen, Dongdong
Gao, Mei
Liu, Mengchen
Yuan, Junsong
Qiao, Chunming
author_facet Feng, Xuelu
Li, Yunsheng
Chen, Dongdong
Gao, Mei
Liu, Mengchen
Yuan, Junsong
Qiao, Chunming
contents Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their deployment faces significant challenges, including slow inference, high computational cost, and impracticality for on-device applications. In contrast, the emergence of small MLLMs, exemplified by the LLava-series models and Phi-3-Vision, offers promising alternatives with faster inference, reduced deployment costs, and the ability to handle domain-specific scenarios. Despite their growing presence, the capability boundaries between large and small MLLMs remain underexplored. In this work, we conduct a systematic and comprehensive evaluation to benchmark both small and large MLLMs, spanning general capabilities such as object recognition, temporal reasoning, and multimodal comprehension, as well as real-world applications in domains like industry and automotive. Our evaluation reveals that small MLLMs can achieve comparable performance to large models in specific scenarios but lag significantly in complex tasks requiring deeper reasoning or nuanced understanding. Furthermore, we identify common failure cases in both small and large MLLMs, highlighting domains where even state-of-the-art models struggle. We hope our findings will guide the research community in pushing the quality boundaries of MLLMs, advancing their usability and effectiveness across diverse applications.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04150
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Large and Small MLLMs
Feng, Xuelu
Li, Yunsheng
Chen, Dongdong
Gao, Mei
Liu, Mengchen
Yuan, Junsong
Qiao, Chunming
Computer Vision and Pattern Recognition
Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their deployment faces significant challenges, including slow inference, high computational cost, and impracticality for on-device applications. In contrast, the emergence of small MLLMs, exemplified by the LLava-series models and Phi-3-Vision, offers promising alternatives with faster inference, reduced deployment costs, and the ability to handle domain-specific scenarios. Despite their growing presence, the capability boundaries between large and small MLLMs remain underexplored. In this work, we conduct a systematic and comprehensive evaluation to benchmark both small and large MLLMs, spanning general capabilities such as object recognition, temporal reasoning, and multimodal comprehension, as well as real-world applications in domains like industry and automotive. Our evaluation reveals that small MLLMs can achieve comparable performance to large models in specific scenarios but lag significantly in complex tasks requiring deeper reasoning or nuanced understanding. Furthermore, we identify common failure cases in both small and large MLLMs, highlighting domains where even state-of-the-art models struggle. We hope our findings will guide the research community in pushing the quality boundaries of MLLMs, advancing their usability and effectiveness across diverse applications.
title Benchmarking Large and Small MLLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04150