Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Yuqi, Shi, Yang, Zhang, Zhuoran, Wang, Qixun, Bai, Xuehai, Ding, Yue, Chen, Ruizhe, Zeng, Bohan, Chen, Xinlong, Zhu, Xuanyu, Li, Bozhou, Wang, Yuran, Dai, Yifan, Tong, Chengzhuo, Liu, Xinyu, Ji, Yiyan, Wei, Yujie, Dong, Yuhao, Yan, Shilin, Wang, Fengxiang, Zhang, Yi-Fan, Wang, Haotian, Zhang, Yuanxing, Wan, Pengfei
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911696312860672
author Tang, Yuqi
Shi, Yang
Zhang, Zhuoran
Wang, Qixun
Bai, Xuehai
Ding, Yue
Chen, Ruizhe
Zeng, Bohan
Chen, Xinlong
Zhu, Xuanyu
Li, Bozhou
Wang, Yuran
Dai, Yifan
Tong, Chengzhuo
Liu, Xinyu
Ji, Yiyan
Wei, Yujie
Dong, Yuhao
Yan, Shilin
Wang, Fengxiang
Zhang, Yi-Fan
Wang, Haotian
Zhang, Yuanxing
Wan, Pengfei
author_facet Tang, Yuqi
Shi, Yang
Zhang, Zhuoran
Wang, Qixun
Bai, Xuehai
Ding, Yue
Chen, Ruizhe
Zeng, Bohan
Chen, Xinlong
Zhu, Xuanyu
Li, Bozhou
Wang, Yuran
Dai, Yifan
Tong, Chengzhuo
Liu, Xinyu
Ji, Yiyan
Wei, Yujie
Dong, Yuhao
Yan, Shilin
Wang, Fengxiang
Zhang, Yi-Fan
Wang, Haotian
Zhang, Yuanxing
Wan, Pengfei
contents Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While Multimodal Large Language Models (MLLMs) show strong visual understanding capabilities, their ability to perceive and reason about such artifacts remains unclear. Existing benchmarks often lack systematic evaluation of artifact-aware perception and fine-grained diagnostic reasoning, especially across diverse AI-generated video domains beyond photorealistic content. To address this gap, we introduce Artifact-Bench, a comprehensive benchmark for evaluating MLLMs on AI-generated video artifact detection and analysis. We first establish a three-level hierarchical taxonomy of realism artifacts, covering photorealistic, animated, and CG-style videos. Based on this taxonomy, Artifact-Bench defines three complementary tasks: real vs. AI-generated video classification, pairwise realism comparison, and fine-grained artifact identification. Experiments on 19 leading MLLMs reveal substantial limitations in artifact perception and reasoning, with many models approaching random or even below-random performance in challenging settings. We further observe significant misalignment between MLLM judgments and human perceptual preferences, highlighting their limited reliability as general evaluators for AI-generated video realism.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18984
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Tang, Yuqi
Shi, Yang
Zhang, Zhuoran
Wang, Qixun
Bai, Xuehai
Ding, Yue
Chen, Ruizhe
Zeng, Bohan
Chen, Xinlong
Zhu, Xuanyu
Li, Bozhou
Wang, Yuran
Dai, Yifan
Tong, Chengzhuo
Liu, Xinyu
Ji, Yiyan
Wei, Yujie
Dong, Yuhao
Yan, Shilin
Wang, Fengxiang
Zhang, Yi-Fan
Wang, Haotian
Zhang, Yuanxing
Wan, Pengfei
Computer Vision and Pattern Recognition
Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While Multimodal Large Language Models (MLLMs) show strong visual understanding capabilities, their ability to perceive and reason about such artifacts remains unclear. Existing benchmarks often lack systematic evaluation of artifact-aware perception and fine-grained diagnostic reasoning, especially across diverse AI-generated video domains beyond photorealistic content. To address this gap, we introduce Artifact-Bench, a comprehensive benchmark for evaluating MLLMs on AI-generated video artifact detection and analysis. We first establish a three-level hierarchical taxonomy of realism artifacts, covering photorealistic, animated, and CG-style videos. Based on this taxonomy, Artifact-Bench defines three complementary tasks: real vs. AI-generated video classification, pairwise realism comparison, and fine-grained artifact identification. Experiments on 19 leading MLLMs reveal substantial limitations in artifact perception and reasoning, with many models approaching random or even below-random performance in challenging settings. We further observe significant misalignment between MLLM judgments and human perceptual preferences, highlighting their limited reliability as general evaluators for AI-generated video realism.
title Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18984