ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Burgess, James, Abdal, Rameen, Stoddart, Dan, Tulyakov, Sergey, Yeung-Levy, Serena, Wang, Kuan-Chieh Jackson
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908825519390720
author Burgess, James
Abdal, Rameen
Stoddart, Dan
Tulyakov, Sergey
Yeung-Levy, Serena
Wang, Kuan-Chieh Jackson
author_facet Burgess, James
Abdal, Rameen
Stoddart, Dan
Tulyakov, Sergey
Yeung-Levy, Serena
Wang, Kuan-Chieh Jackson
contents Modern image generators produce strikingly realistic images, where only artifacts like distorted hands or warped objects reveal their synthetic origin. Detecting these artifacts is essential: without detection, we cannot benchmark generators or train reward models to improve them. Current detectors fine-tune VLMs on tens of thousands of labeled images, but this is expensive to repeat whenever generators evolve or new artifact types emerge. We show that pretrained VLMs already encode the knowledge needed to detect artifacts - with the right scaffolding, this capability can be unlocked using only a few hundred labeled examples per artifact category. Our system, ArtifactLens, achieves state-of-the-art on five human artifact benchmarks (the first evaluation across multiple datasets) while requiring orders of magnitude less labeled data. The scaffolding consists of a multi-component architecture with in-context learning and text instruction optimization, with novel improvements to each. Our methods generalize to other artifact types - object morphology, animal anatomy, and entity interactions - and to the distinct task of AIGC detection.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09475
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
Burgess, James
Abdal, Rameen
Stoddart, Dan
Tulyakov, Sergey
Yeung-Levy, Serena
Wang, Kuan-Chieh Jackson
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Modern image generators produce strikingly realistic images, where only artifacts like distorted hands or warped objects reveal their synthetic origin. Detecting these artifacts is essential: without detection, we cannot benchmark generators or train reward models to improve them. Current detectors fine-tune VLMs on tens of thousands of labeled images, but this is expensive to repeat whenever generators evolve or new artifact types emerge. We show that pretrained VLMs already encode the knowledge needed to detect artifacts - with the right scaffolding, this capability can be unlocked using only a few hundred labeled examples per artifact category. Our system, ArtifactLens, achieves state-of-the-art on five human artifact benchmarks (the first evaluation across multiple datasets) while requiring orders of magnitude less labeled data. The scaffolding consists of a multi-component architecture with in-context learning and text instruction optimization, with novel improvements to each. Our methods generalize to other artifact types - object morphology, animal anatomy, and entity interactions - and to the distinct task of AIGC detection.
title ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.09475