AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Junyang, Wang, Yuhang, Xu, Guohai, Zhang, Jing, Gu, Yukai, Jia, Haitao, Wang, Jiaqi, Xu, Haiyang, Yan, Ming, Zhang, Ji, Sang, Jitao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917595919155200
author Wang, Junyang
Wang, Yuhang
Xu, Guohai
Zhang, Jing
Gu, Yukai
Jia, Haitao
Wang, Jiaqi
Xu, Haiyang
Yan, Ming
Zhang, Ji
Sang, Jitao
author_facet Wang, Junyang
Wang, Yuhang
Xu, Guohai
Zhang, Jing
Gu, Yukai
Jia, Haitao
Wang, Jiaqi
Xu, Haiyang
Yan, Ming
Zhang, Ji
Sang, Jitao
contents Despite making significant progress in multi-modal tasks, current Multi-modal Large Language Models (MLLMs) encounter the significant challenge of hallucinations, which may lead to harmful consequences. Therefore, evaluating MLLMs' hallucinations is becoming increasingly important in model improvement and practical application deployment. Previous works are limited in high evaluation costs (e.g., relying on humans or advanced LLMs) and insufficient evaluation dimensions (e.g., types of tasks and hallucinations). In this paper, we propose an LLM-free multi-dimensional benchmark AMBER, which can be used to evaluate both generative task and discriminative task including existence, attribute and relation hallucination. Based on AMBER, we design a low-cost and efficient evaluation pipeline. Additionally, we conduct a comprehensive evaluation and detailed analysis of mainstream MLLMs including GPT-4V(ision), and also give guideline suggestions for mitigating hallucinations. The data and code of AMBER are available at https://github.com/junyangwang0410/AMBER.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07397
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
Wang, Junyang
Wang, Yuhang
Xu, Guohai
Zhang, Jing
Gu, Yukai
Jia, Haitao
Wang, Jiaqi
Xu, Haiyang
Yan, Ming
Zhang, Ji
Sang, Jitao
Computation and Language
Computer Vision and Pattern Recognition
Despite making significant progress in multi-modal tasks, current Multi-modal Large Language Models (MLLMs) encounter the significant challenge of hallucinations, which may lead to harmful consequences. Therefore, evaluating MLLMs' hallucinations is becoming increasingly important in model improvement and practical application deployment. Previous works are limited in high evaluation costs (e.g., relying on humans or advanced LLMs) and insufficient evaluation dimensions (e.g., types of tasks and hallucinations). In this paper, we propose an LLM-free multi-dimensional benchmark AMBER, which can be used to evaluate both generative task and discriminative task including existence, attribute and relation hallucination. Based on AMBER, we design a low-cost and efficient evaluation pipeline. Additionally, we conduct a comprehensive evaluation and detailed analysis of mainstream MLLMs including GPT-4V(ision), and also give guideline suggestions for mitigating hallucinations. The data and code of AMBER are available at https://github.com/junyangwang0410/AMBER.
title AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.07397