The Visual Counter Turing Test (VCT2): A Benchmark for Evaluating AI-Generated Image Detection and the Visual AI Index (VAI)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Imanpour, Nasrin, Borah, Abhilekh, Bajpai, Shashwat, Ghosh, Subhankar, Sankepally, Sainath Reddy, Abdullah, Hasnat Md, Kosaraju, Nishoak, Dixit, Shreyas, Aziz, Ashhar, Biswas, Shwetangshu, Jain, Vinija, Chadha, Aman, Wang, Song, Sheth, Amit, Das, Amitava
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915612334227456
author Imanpour, Nasrin
Borah, Abhilekh
Bajpai, Shashwat
Ghosh, Subhankar
Sankepally, Sainath Reddy
Abdullah, Hasnat Md
Kosaraju, Nishoak
Dixit, Shreyas
Aziz, Ashhar
Biswas, Shwetangshu
Jain, Vinija
Chadha, Aman
Wang, Song
Sheth, Amit
Das, Amitava
author_facet Imanpour, Nasrin
Borah, Abhilekh
Bajpai, Shashwat
Ghosh, Subhankar
Sankepally, Sainath Reddy
Abdullah, Hasnat Md
Kosaraju, Nishoak
Dixit, Shreyas
Aziz, Ashhar
Biswas, Shwetangshu
Jain, Vinija
Chadha, Aman
Wang, Song
Sheth, Amit
Das, Amitava
contents The rapid progress and widespread availability of text-to-image (T2I) generative models have heightened concerns about the misuse of AI-generated visuals, particularly in the context of misinformation campaigns. Existing AI-generated image detection (AGID) methods often overfit to known generators and falter on outputs from newer or unseen models. We introduce the Visual Counter Turing Test (VCT2), a comprehensive benchmark of 166,000 images, comprising both real and synthetic prompt-image pairs produced by six state-of-the-art T2I systems: Stable Diffusion 2.1, SDXL, SD3 Medium, SD3.5 Large, DALL.E 3, and Midjourney 6. We curate two distinct subsets: COCOAI, featuring structured captions from MS COCO, and TwitterAI, containing narrative-style tweets from The New York Times. Under a unified zero-shot evaluation, we benchmark 17 leading AGID models and observe alarmingly low detection accuracy, 58% on COCOAI and 58.34% on TwitterAI. To transcend binary classification, we propose the Visual AI Index (VAI), an interpretable, prompt-agnostic realism metric based on twelve low-level visual features, enabling us to quantify and rank the perceptual quality of generated outputs with greater nuance. Correlation analysis reveals a moderate inverse relationship between VAI and detection accuracy: Pearson of -0.532 on COCOAI and -0.503 on TwitterAI, suggesting that more visually realistic images tend to be harder to detect, a trend observed consistently across generators. We release COCOAI, TwitterAI, and all codes to catalyze future advances in generalized AGID and perceptual realism assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16754
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Visual Counter Turing Test (VCT2): A Benchmark for Evaluating AI-Generated Image Detection and the Visual AI Index (VAI)
Imanpour, Nasrin
Borah, Abhilekh
Bajpai, Shashwat
Ghosh, Subhankar
Sankepally, Sainath Reddy
Abdullah, Hasnat Md
Kosaraju, Nishoak
Dixit, Shreyas
Aziz, Ashhar
Biswas, Shwetangshu
Jain, Vinija
Chadha, Aman
Wang, Song
Sheth, Amit
Das, Amitava
Computer Vision and Pattern Recognition
Artificial Intelligence
The rapid progress and widespread availability of text-to-image (T2I) generative models have heightened concerns about the misuse of AI-generated visuals, particularly in the context of misinformation campaigns. Existing AI-generated image detection (AGID) methods often overfit to known generators and falter on outputs from newer or unseen models. We introduce the Visual Counter Turing Test (VCT2), a comprehensive benchmark of 166,000 images, comprising both real and synthetic prompt-image pairs produced by six state-of-the-art T2I systems: Stable Diffusion 2.1, SDXL, SD3 Medium, SD3.5 Large, DALL.E 3, and Midjourney 6. We curate two distinct subsets: COCOAI, featuring structured captions from MS COCO, and TwitterAI, containing narrative-style tweets from The New York Times. Under a unified zero-shot evaluation, we benchmark 17 leading AGID models and observe alarmingly low detection accuracy, 58% on COCOAI and 58.34% on TwitterAI. To transcend binary classification, we propose the Visual AI Index (VAI), an interpretable, prompt-agnostic realism metric based on twelve low-level visual features, enabling us to quantify and rank the perceptual quality of generated outputs with greater nuance. Correlation analysis reveals a moderate inverse relationship between VAI and detection accuracy: Pearson of -0.532 on COCOAI and -0.503 on TwitterAI, suggesting that more visually realistic images tend to be harder to detect, a trend observed consistently across generators. We release COCOAI, TwitterAI, and all codes to catalyze future advances in generalized AGID and perceptual realism assessment.
title The Visual Counter Turing Test (VCT2): A Benchmark for Evaluating AI-Generated Image Detection and the Visual AI Index (VAI)
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.16754