Detecting Text Manipulation in Images using Vision Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914033569890304 |
|---|---|
| author | Vidit, Vidit Korshunov, Pavel Mohammadi, Amir Ecabert, Christophe Kotwal, Ketan Marcel, Sébastien |
| author_facet | Vidit, Vidit Korshunov, Pavel Mohammadi, Amir Ecabert, Christophe Kotwal, Ketan Marcel, Sébastien |
| contents | Recent works have shown the effectiveness of Large Vision Language Models (VLMs or LVLMs) in image manipulation detection. However, text manipulation detection is largely missing in these studies. We bridge this knowledge gap by analyzing closed- and open-source VLMs on different text manipulation datasets. Our results suggest that open-source models are getting closer, but still behind closed-source ones like GPT- 4o. Additionally, we benchmark image manipulation detection-specific VLMs for text manipulation detection and show that they suffer from the generalization problem. We benchmark VLMs for manipulations done on in-the-wild scene texts and on fantasy ID cards, where the latter mimic a challenging real-world misuse. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_10278 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Detecting Text Manipulation in Images using Vision Language Models Vidit, Vidit Korshunov, Pavel Mohammadi, Amir Ecabert, Christophe Kotwal, Ketan Marcel, Sébastien Computer Vision and Pattern Recognition Recent works have shown the effectiveness of Large Vision Language Models (VLMs or LVLMs) in image manipulation detection. However, text manipulation detection is largely missing in these studies. We bridge this knowledge gap by analyzing closed- and open-source VLMs on different text manipulation datasets. Our results suggest that open-source models are getting closer, but still behind closed-source ones like GPT- 4o. Additionally, we benchmark image manipulation detection-specific VLMs for text manipulation detection and show that they suffer from the generalization problem. We benchmark VLMs for manipulations done on in-the-wild scene texts and on fantasy ID cards, where the latter mimic a challenging real-world misuse. |
| title | Detecting Text Manipulation in Images using Vision Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.10278 |