FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Zheqi, Liu, Yesheng, Zheng, Jing-shu, Li, Xuejing, Yao, Jin-Ge, Qin, Bowen, Xuan, Richeng, Yang, Xi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916867777495040
author He, Zheqi
Liu, Yesheng
Zheng, Jing-shu
Li, Xuejing
Yao, Jin-Ge
Qin, Bowen
Xuan, Richeng
Yang, Xi
author_facet He, Zheqi
Liu, Yesheng
Zheng, Jing-shu
Li, Xuejing
Yao, Jin-Ge
Qin, Bowen
Xuan, Richeng
Yang, Xi
contents We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answering, text-to-image/video generation, and image-text retrieval. We decouple model inference from evaluation through an independent evaluation service, thus enabling flexible resource allocation and seamless integration of new tasks and models. Moreover, FlagEvalMM utilizes advanced inference acceleration tools (e.g., vLLM, SGLang) and asynchronous data loading to significantly enhance evaluation efficiency. Extensive experiments show that FlagEvalMM offers accurate and efficient insights into model strengths and limitations, making it a valuable tool for advancing multimodal research. The framework is publicly accessible at https://github.com/flageval-baai/FlagEvalMM.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
He, Zheqi
Liu, Yesheng
Zheng, Jing-shu
Li, Xuejing
Yao, Jin-Ge
Qin, Bowen
Xuan, Richeng
Yang, Xi
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answering, text-to-image/video generation, and image-text retrieval. We decouple model inference from evaluation through an independent evaluation service, thus enabling flexible resource allocation and seamless integration of new tasks and models. Moreover, FlagEvalMM utilizes advanced inference acceleration tools (e.g., vLLM, SGLang) and asynchronous data loading to significantly enhance evaluation efficiency. Extensive experiments show that FlagEvalMM offers accurate and efficient insights into model strengths and limitations, making it a valuable tool for advancing multimodal research. The framework is publicly accessible at https://github.com/flageval-baai/FlagEvalMM.
title FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.09081