Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Waseda, Futa, Yamabe, Shojiro, Shiono, Daiki, Sasaki, Kento, Takahashi, Tsubasa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
by: Yamabe, Shojiro, et al.
Published: (2025)
by: Yamabe, Shojiro, et al.
Published: (2025)
Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
by: Takahashi, Tsubasa, et al.
Published: (2025)
by: Takahashi, Tsubasa, et al.
Published: (2025)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
by: Waseda, Futa, et al.
Published: (2025)
by: Waseda, Futa, et al.
Published: (2025)
MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models
by: Yamabe, Shojiro, et al.
Published: (2024)
by: Yamabe, Shojiro, et al.
Published: (2024)
STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes
by: Ishihara, Keishi, et al.
Published: (2025)
by: Ishihara, Keishi, et al.
Published: (2025)
Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
by: Waseda, Futa, et al.
Published: (2024)
by: Waseda, Futa, et al.
Published: (2024)
Instruction-Following Evaluation of Large Vision-Language Models
by: Shiono, Daiki, et al.
Published: (2025)
by: Shiono, Daiki, et al.
Published: (2025)
Defending Against Physical Adversarial Patch Attacks on Infrared Human Detection
by: Strack, Lukas, et al.
Published: (2023)
by: Strack, Lukas, et al.
Published: (2023)
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
by: Yokoi, Shingo, et al.
Published: (2025)
by: Yokoi, Shingo, et al.
Published: (2025)
One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
by: Miwa, Keita, et al.
Published: (2025)
by: Miwa, Keita, et al.
Published: (2025)
Uncolorable Examples: Preventing Unauthorized AI Colorization via Perception-Aware Chroma-Restrictive Perturbation
by: Nii, Yuki, et al.
Published: (2025)
by: Nii, Yuki, et al.
Published: (2025)
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Typographic Text Generation with Off-the-Shelf Diffusion Model
by: Peong, KhayTze, et al.
Published: (2024)
by: Peong, KhayTze, et al.
Published: (2024)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
by: Inoue, Yuichi, et al.
Published: (2024)
by: Inoue, Yuichi, et al.
Published: (2024)
Automatic Text Box Placement for Supporting Typographic Design
by: Muraoka, Jun, et al.
Published: (2025)
by: Muraoka, Jun, et al.
Published: (2025)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
by: Zhu, Jiayi, et al.
Published: (2025)
by: Zhu, Jiayi, et al.
Published: (2025)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
by: Hufe, Lorenz, et al.
Published: (2025)
by: Hufe, Lorenz, et al.
Published: (2025)
Benchmarking Large Language Models for Handwritten Text Recognition
by: Crosilla, Giorgia, et al.
Published: (2025)
by: Crosilla, Giorgia, et al.
Published: (2025)
One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models
by: Zhao, Jiale, et al.
Published: (2025)
by: Zhao, Jiale, et al.
Published: (2025)
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
by: Chen, Tianle, et al.
Published: (2026)
by: Chen, Tianle, et al.
Published: (2026)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
by: Westerhoff, Justus, et al.
Published: (2025)
by: Westerhoff, Justus, et al.
Published: (2025)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
by: Li, Yueyan, et al.
Published: (2025)
by: Li, Yueyan, et al.
Published: (2025)
Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
by: Xie, Peng, et al.
Published: (2024)
by: Xie, Peng, et al.
Published: (2024)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
by: Xu, Jingning, et al.
Published: (2026)
by: Xu, Jingning, et al.
Published: (2026)
RobustNeRF: Ignoring Distractors with Robust Losses
by: Sabour, Sara, et al.
Published: (2023)
by: Sabour, Sara, et al.
Published: (2023)
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
by: Arai, Hidehisa, et al.
Published: (2024)
by: Arai, Hidehisa, et al.
Published: (2024)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
PHyCLIP: $\ell_1$-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
by: Yoshikawa, Daiki, et al.
Published: (2025)
by: Yoshikawa, Daiki, et al.
Published: (2025)
LaVPR: Benchmarking Language and Vision for Place Recognition
by: Idan, Ofer, et al.
Published: (2026)
by: Idan, Ofer, et al.
Published: (2026)
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
by: Nagaonkar, Sankalp, et al.
Published: (2025)
by: Nagaonkar, Sankalp, et al.
Published: (2025)
Unified Prompt Attack Against Text-to-Image Generation Models
by: Peng, Duo, et al.
Published: (2025)
by: Peng, Duo, et al.
Published: (2025)
Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality
by: Xiu, Yanming, et al.
Published: (2026)
by: Xiu, Yanming, et al.
Published: (2026)
Similar Items
-
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
by: Yamabe, Shojiro, et al.
Published: (2025) -
Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
by: Takahashi, Tsubasa, et al.
Published: (2025) -
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
by: Waseda, Futa, et al.
Published: (2025) -
MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models
by: Yamabe, Shojiro, et al.
Published: (2024) -
STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes
by: Ishihara, Keishi, et al.
Published: (2025)