Detecting Text Manipulation in Images using Vision Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vidit, Vidit, Korshunov, Pavel, Mohammadi, Amir, Ecabert, Christophe, Kotwal, Ketan, Marcel, Sébastien
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914033569890304
author Vidit, Vidit
Korshunov, Pavel
Mohammadi, Amir
Ecabert, Christophe
Kotwal, Ketan
Marcel, Sébastien
author_facet Vidit, Vidit
Korshunov, Pavel
Mohammadi, Amir
Ecabert, Christophe
Kotwal, Ketan
Marcel, Sébastien
contents Recent works have shown the effectiveness of Large Vision Language Models (VLMs or LVLMs) in image manipulation detection. However, text manipulation detection is largely missing in these studies. We bridge this knowledge gap by analyzing closed- and open-source VLMs on different text manipulation datasets. Our results suggest that open-source models are getting closer, but still behind closed-source ones like GPT- 4o. Additionally, we benchmark image manipulation detection-specific VLMs for text manipulation detection and show that they suffer from the generalization problem. We benchmark VLMs for manipulations done on in-the-wild scene texts and on fantasy ID cards, where the latter mimic a challenging real-world misuse.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting Text Manipulation in Images using Vision Language Models
Vidit, Vidit
Korshunov, Pavel
Mohammadi, Amir
Ecabert, Christophe
Kotwal, Ketan
Marcel, Sébastien
Computer Vision and Pattern Recognition
Recent works have shown the effectiveness of Large Vision Language Models (VLMs or LVLMs) in image manipulation detection. However, text manipulation detection is largely missing in these studies. We bridge this knowledge gap by analyzing closed- and open-source VLMs on different text manipulation datasets. Our results suggest that open-source models are getting closer, but still behind closed-source ones like GPT- 4o. Additionally, we benchmark image manipulation detection-specific VLMs for text manipulation detection and show that they suffer from the generalization problem. We benchmark VLMs for manipulations done on in-the-wild scene texts and on fantasy ID cards, where the latter mimic a challenging real-world misuse.
title Detecting Text Manipulation in Images using Vision Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.10278