Image Realness Assessment and Localization with Multimodal Features

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kaushik, Lovish, Biswas, Agnij, Paul, Somdyuti
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908542902992896
author Kaushik, Lovish
Biswas, Agnij
Paul, Somdyuti
author_facet Kaushik, Lovish
Biswas, Agnij
Paul, Somdyuti
contents A reliable method of quantifying the perceptual realness of AI-generated images and identifying visually inconsistent regions is crucial for practical use of AI-generated images and for improving photorealism of generative AI via realness feedback during training. This paper introduces a framework that accomplishes both overall objective realness assessment and local inconsistency identification of AI-generated images using textual descriptions of visual inconsistencies generated by vision-language models trained on large datasets that serve as reliable substitutes for human annotations. Our results demonstrate that the proposed multimodal approach improves objective realness prediction performance and produces dense realness maps that effectively distinguish between realistic and unrealistic spatial regions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13289
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Image Realness Assessment and Localization with Multimodal Features
Kaushik, Lovish
Biswas, Agnij
Paul, Somdyuti
Computer Vision and Pattern Recognition
Image and Video Processing
A reliable method of quantifying the perceptual realness of AI-generated images and identifying visually inconsistent regions is crucial for practical use of AI-generated images and for improving photorealism of generative AI via realness feedback during training. This paper introduces a framework that accomplishes both overall objective realness assessment and local inconsistency identification of AI-generated images using textual descriptions of visual inconsistencies generated by vision-language models trained on large datasets that serve as reliable substitutes for human annotations. Our results demonstrate that the proposed multimodal approach improves objective realness prediction performance and produces dense realness maps that effectively distinguish between realistic and unrealistic spatial regions.
title Image Realness Assessment and Localization with Multimodal Features
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2509.13289