Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Yuqi, Shi, Yang, Zhang, Zhuoran, Wang, Qixun, Bai, Xuehai, Ding, Yue, Chen, Ruizhe, Zeng, Bohan, Chen, Xinlong, Zhu, Xuanyu, Li, Bozhou, Wang, Yuran, Dai, Yifan, Tong, Chengzhuo, Liu, Xinyu, Ji, Yiyan, Wei, Yujie, Dong, Yuhao, Yan, Shilin, Wang, Fengxiang, Zhang, Yi-Fan, Wang, Haotian, Zhang, Yuanxing, Wan, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
by: Liu, Tengfei, et al.
Published: (2026)
by: Liu, Tengfei, et al.
Published: (2026)
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
by: Zhang, Chenchen, et al.
Published: (2025)
by: Zhang, Chenchen, et al.
Published: (2025)
Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
by: Li, Bozhou, et al.
Published: (2026)
by: Li, Bozhou, et al.
Published: (2026)
OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
by: Ding, Yue, et al.
Published: (2026)
by: Ding, Yue, et al.
Published: (2026)
Detecting Human Artifacts from Text-to-Image Models
by: Wang, Kaihong, et al.
Published: (2024)
by: Wang, Kaihong, et al.
Published: (2024)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
GenShield: Unified Detection and Artifact Correction for AI-Generated Images
by: Xu, Zhipei, et al.
Published: (2026)
by: Xu, Zhipei, et al.
Published: (2026)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
Agent-Based Software Artifact Evaluation
by: Wu, Zhaonan, et al.
Published: (2026)
by: Wu, Zhaonan, et al.
Published: (2026)
Zero-Shot Artifact2Artifact: Self-incentive artifact removal for photoacoustic imaging without any data
by: Li, Shuang, et al.
Published: (2024)
by: Li, Shuang, et al.
Published: (2024)
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
by: Zhang, Zhuoran, et al.
Published: (2025)
by: Zhang, Zhuoran, et al.
Published: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
by: Chen, Xinlong, et al.
Published: (2026)
by: Chen, Xinlong, et al.
Published: (2026)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
Figures as Interfaces: Toward LLM-Native Artifacts for Scientific Discovery
by: Wang, Yifang, et al.
Published: (2026)
by: Wang, Yifang, et al.
Published: (2026)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
by: Chen, Zhihong, et al.
Published: (2025)
by: Chen, Zhihong, et al.
Published: (2025)
Artifacts of the ADS Bug-Fix Pattern Study
by: Chen, Yuntianyi, et al.
Published: (2025)
by: Chen, Yuntianyi, et al.
Published: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
TestLoc Artifact
by: Chen, Yang
Published: (2026)
by: Chen, Yang
Published: (2026)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
by: Li, Caorui, et al.
Published: (2025)
by: Li, Caorui, et al.
Published: (2025)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
by: Wang, Juan, et al.
Published: (2026)
by: Wang, Juan, et al.
Published: (2026)
Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
by: Zeng, Bohan, et al.
Published: (2026)
by: Zeng, Bohan, et al.
Published: (2026)
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models
by: Wang, Xinliang, et al.
Published: (2026)
by: Wang, Xinliang, et al.
Published: (2026)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image Detection
by: Wang, Wenbin, et al.
Published: (2026)
by: Wang, Wenbin, et al.
Published: (2026)
Usenix security 25' cycle1-190-ALERT-Artifact-Evaluation
by: Wang, Longxiang, et al.
Published: (2025)
by: Wang, Longxiang, et al.
Published: (2025)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
Rainbow Artifacts from Electromagnetic Signal Injection Attacks on Image Sensors
by: Zhang, Youqian, et al.
Published: (2025)
by: Zhang, Youqian, et al.
Published: (2025)
Noise-Space Attribution and Control of Chunk-Boundary Artifact
by: Wang, Rui
Published: (2026)
by: Wang, Rui
Published: (2026)
LEGO: HPCA2026 Artifact Evaluation
by: Han, Zhao, et al.
Published: (2025)
by: Han, Zhao, et al.
Published: (2025)
A Deep Learning-Based Method for Metal Artifact-Resistant Syn-MP-RAGE Contrast Synthesis
by: Zeng, Ziyi, et al.
Published: (2024)
by: Zeng, Ziyi, et al.
Published: (2024)
Diffusion Model Regularized Implicit Neural Representation for CT Metal Artifact Reduction
by: Wen, Jie, et al.
Published: (2025)
by: Wen, Jie, et al.
Published: (2025)
VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining
by: Zhu, Xuanyu, et al.
Published: (2026)
by: Zhu, Xuanyu, et al.
Published: (2026)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
by: Meng, Jiahao, et al.
Published: (2026)
by: Meng, Jiahao, et al.
Published: (2026)
Similar Items
-
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
by: Liu, Tengfei, et al.
Published: (2026) -
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
by: Shi, Yang, et al.
Published: (2025) -
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025) -
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026) -
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)