Detecting Text Manipulation in Images using Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Vidit, Vidit, Korshunov, Pavel, Mohammadi, Amir, Ecabert, Christophe, Kotwal, Ketan, Marcel, Sébastien |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data
by: Korshunov, Pavel, et al.
Published: (2025)
by: Korshunov, Pavel, et al.
Published: (2025)
FantasyID: A dataset for detecting digital manipulations of ID-documents
by: Korshunov, Pavel, et al.
Published: (2025)
by: Korshunov, Pavel, et al.
Published: (2025)
EdgeFace: Efficient Face Recognition Model for Edge Devices
by: George, Anjith, et al.
Published: (2023)
by: George, Anjith, et al.
Published: (2023)
Producing Histopathology Phantom Images using Generative Adversarial Networks to improve Tumor Detection
by: Gautam, Vidit
Published: (2022)
by: Gautam, Vidit
Published: (2022)
Review of Demographic Fairness in Face Recognition
by: Kotwal, Ketan, et al.
Published: (2025)
by: Kotwal, Ketan, et al.
Published: (2025)
Addressing the Elephant in the Room: Robust Animal Re-Identification with Unsupervised Part-Based Feature Alignment
by: Yu, Yingxue, et al.
Published: (2024)
by: Yu, Yingxue, et al.
Published: (2024)
Identity-Preserving Aging and De-Aging of Faces in the StyleGAN Latent Space
by: Luevano, Luis S., et al.
Published: (2025)
by: Luevano, Luis S., et al.
Published: (2025)
VRBiom: A New Periocular Dataset for Biometric Applications of HMD
by: Kotwal, Ketan, et al.
Published: (2024)
by: Kotwal, Ketan, et al.
Published: (2024)
Latent Enhancing AutoEncoder for Occluded Image Classification
by: Kotwal, Ketan, et al.
Published: (2024)
by: Kotwal, Ketan, et al.
Published: (2024)
Score Normalization for Demographic Fairness in Face Recognition
by: Linghu, Yu, et al.
Published: (2024)
by: Linghu, Yu, et al.
Published: (2024)
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
VASE: Object-Centric Appearance and Shape Manipulation of Real Videos
by: Peruzzo, Elia, et al.
Published: (2024)
by: Peruzzo, Elia, et al.
Published: (2024)
Predicting Performance of Object Detection Models in Electron Microscopy Using Random Forests
by: Li, Ni, et al.
Published: (2025)
by: Li, Ni, et al.
Published: (2025)
$\textit{sweet}$- An Open Source Modular Platform for Contactless Hand Vascular Biometric Experiments
by: Geissbühler, David, et al.
Published: (2024)
by: Geissbühler, David, et al.
Published: (2024)
Federated Action Recognition for Smart Worker Assistance Using FastPose
by: Hegiste, Vinit, et al.
Published: (2025)
by: Hegiste, Vinit, et al.
Published: (2025)
Wonderland: Navigating 3D Scenes from a Single Image
by: Liang, Hanwen, et al.
Published: (2024)
by: Liang, Hanwen, et al.
Published: (2024)
Hand Shape and Gesture Recognition using Multiscale Template Matching, Background Subtraction and Binary Image Analysis
by: Saichandran, Ketan Suhaas
Published: (2024)
by: Saichandran, Ketan Suhaas
Published: (2024)
Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach
by: Chen, Yunuo, et al.
Published: (2025)
by: Chen, Yunuo, et al.
Published: (2025)
EdgeDoc: Hybrid CNN-Transformer Model for Accurate Forgery Detection and Localization in ID Documents
by: George, Anjith, et al.
Published: (2025)
by: George, Anjith, et al.
Published: (2025)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
by: Li, Yueyan, et al.
Published: (2025)
by: Li, Yueyan, et al.
Published: (2025)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023)
by: Goel, Vidit, et al.
Published: (2023)
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
by: Johnson, Emily, et al.
Published: (2025)
by: Johnson, Emily, et al.
Published: (2025)
Tag2Text: Guiding Vision-Language Model via Image Tagging
by: Huang, Xinyu, et al.
Published: (2023)
by: Huang, Xinyu, et al.
Published: (2023)
Leveraging Vision-Language Models to Detect Attention in Educational Videos
by: Becquet, Gabriel, et al.
Published: (2026)
by: Becquet, Gabriel, et al.
Published: (2026)
AugGen: Synthetic Augmentation using Diffusion Models Can Improve Recognition
by: Rahimi, Parsa, et al.
Published: (2025)
by: Rahimi, Parsa, et al.
Published: (2025)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)
by: Hayes, Kevin David, et al.
Published: (2025)
Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition
by: Shahreza, Hatef Otroshi, et al.
Published: (2026)
by: Shahreza, Hatef Otroshi, et al.
Published: (2026)
FIDAVL: Fake Image Detection and Attribution using Vision-Language Model
by: Keita, Mamadou, et al.
Published: (2024)
by: Keita, Mamadou, et al.
Published: (2024)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models
by: Zhao, Rui, et al.
Published: (2024)
by: Zhao, Rui, et al.
Published: (2024)
Attention! Your Vision Language Model Could Be Maliciously Manipulated
by: Wang, Xiaosen, et al.
Published: (2025)
by: Wang, Xiaosen, et al.
Published: (2025)
Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
by: Pani, Anupam, et al.
Published: (2026)
by: Pani, Anupam, et al.
Published: (2026)
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025)
by: Park, Jaeyoo, et al.
Published: (2025)
Online Gaussian Test-Time Adaptation of Vision-Language Models
by: Fuchs, Clément, et al.
Published: (2025)
by: Fuchs, Clément, et al.
Published: (2025)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
by: Jia, Mengzhao, et al.
Published: (2024)
by: Jia, Mengzhao, et al.
Published: (2024)
Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models
by: Koma, Arian Komaei, et al.
Published: (2026)
by: Koma, Arian Komaei, et al.
Published: (2026)
SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
Similar Items
-
Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data
by: Korshunov, Pavel, et al.
Published: (2025) -
FantasyID: A dataset for detecting digital manipulations of ID-documents
by: Korshunov, Pavel, et al.
Published: (2025) -
EdgeFace: Efficient Face Recognition Model for Edge Devices
by: George, Anjith, et al.
Published: (2023) -
Producing Histopathology Phantom Images using Generative Adversarial Networks to improve Tumor Detection
by: Gautam, Vidit
Published: (2022) -
Review of Demographic Fairness in Face Recognition
by: Kotwal, Ketan, et al.
Published: (2025)