Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Hao, Xiao, Erjia, Gu, Jindong, Yang, Le, Duan, Jinhao, Zhang, Jize, Cao, Jiahang, Xu, Kaidi, Xu, Renjing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models
by: Cheng, Hao, et al.
Published: (2025)
by: Cheng, Hao, et al.
Published: (2025)
Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Gaining the Sparse Rewards by Exploring Lottery Tickets in Spiking Neural Network
by: Cheng, Hao, et al.
Published: (2023)
by: Cheng, Hao, et al.
Published: (2023)
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
by: Zhu, Jiayi, et al.
Published: (2025)
by: Zhu, Jiayi, et al.
Published: (2025)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
Typographic Text Generation with Off-the-Shelf Diffusion Model
by: Peong, KhayTze, et al.
Published: (2024)
by: Peong, KhayTze, et al.
Published: (2024)
Automatic Text Box Placement for Supporting Typographic Design
by: Muraoka, Jun, et al.
Published: (2025)
by: Muraoka, Jun, et al.
Published: (2025)
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
by: Waseda, Futa, et al.
Published: (2025)
by: Waseda, Futa, et al.
Published: (2025)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
by: Hufe, Lorenz, et al.
Published: (2025)
by: Hufe, Lorenz, et al.
Published: (2025)
Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
by: Gong, Yichen, et al.
Published: (2023)
by: Gong, Yichen, et al.
Published: (2023)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
by: Cheng, Hao, et al.
Published: (2025)
by: Cheng, Hao, et al.
Published: (2025)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
by: Chen, Tianle, et al.
Published: (2026)
by: Chen, Tianle, et al.
Published: (2026)
SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
by: Westerhoff, Justus, et al.
Published: (2025)
by: Westerhoff, Justus, et al.
Published: (2025)
ACT-Diffusion: Efficient Adversarial Consistency Training for One-step Diffusion Models
by: Kong, Fei, et al.
Published: (2023)
by: Kong, Fei, et al.
Published: (2023)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
by: Gan, Esther, et al.
Published: (2024)
by: Gan, Esther, et al.
Published: (2024)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
Adaptive Calibration: A Unified Conversion Framework of Spiking Neural Networks
by: Wang, Ziqing, et al.
Published: (2023)
by: Wang, Ziqing, et al.
Published: (2023)
LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
by: Kong, Fei, et al.
Published: (2025)
by: Kong, Fei, et al.
Published: (2025)
One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
by: Duan, Jinhao, et al.
Published: (2023)
by: Duan, Jinhao, et al.
Published: (2023)
Modality-Composable Diffusion Policy via Inference-Time Distribution-level Composition
by: Cao, Jiahang, et al.
Published: (2025)
by: Cao, Jiahang, et al.
Published: (2025)
Evaluation Metrics for Automated Typographic Poster Generation
by: Rebelo, Sérgio M., et al.
Published: (2024)
by: Rebelo, Sérgio M., et al.
Published: (2024)
Typographic Attacks in a Multi-Image Setting
by: Wang, Xiaomeng, et al.
Published: (2025)
by: Wang, Xiaomeng, et al.
Published: (2025)
Spatial and Typographic Coding in Printed Bibliographical Materials
by: Spencer, H., et al.
Published: (1975)
by: Spencer, H., et al.
Published: (1975)
Joseph Ames's "Typographical Antiquities" and the Antiquarian Tradition
by: Shiner, Elaine
Published: (2013)
by: Shiner, Elaine
Published: (2013)
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
by: Fan, Haozhi, et al.
Published: (2026)
by: Fan, Haozhi, et al.
Published: (2026)
Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors
by: Tsuji, Kohei, et al.
Published: (2025)
by: Tsuji, Kohei, et al.
Published: (2025)
Robust SG-NeRF: Robust Scene Graph Aided Neural Surface Reconstruction
by: Gu, Yi, et al.
Published: (2024)
by: Gu, Yi, et al.
Published: (2024)
Event Masked Autoencoder: Point-wise Action Recognition with Event-Based Cameras
by: Sun, Jingkai, et al.
Published: (2025)
by: Sun, Jingkai, et al.
Published: (2025)
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
by: Li, Yanjie, et al.
Published: (2025)
by: Li, Yanjie, et al.
Published: (2025)
Spiking Diffusion Models
by: Cao, Jiahang, et al.
Published: (2024)
by: Cao, Jiahang, et al.
Published: (2024)
Spiking Neural Network as Adaptive Event Stream Slicer
by: Cao, Jiahang, et al.
Published: (2024)
by: Cao, Jiahang, et al.
Published: (2024)
Adaptive Calibration: A Unified Conversion Framework of Spiking Neural Network
by: Wang, Ziqing, et al.
Published: (2024)
by: Wang, Ziqing, et al.
Published: (2024)
Chasing Day and Night: Towards Robust and Efficient All-Day Object Detection Guided by an Event Camera
by: Cao, Jiahang, et al.
Published: (2023)
by: Cao, Jiahang, et al.
Published: (2023)
Sparse View Distractor-Free Gaussian Splatting
by: Gu, Yi, et al.
Published: (2026)
by: Gu, Yi, et al.
Published: (2026)
Similar Items
-
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024) -
Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models
by: Cheng, Hao, et al.
Published: (2025) -
Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models
by: Cheng, Hao, et al.
Published: (2024) -
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
by: Cheng, Hao, et al.
Published: (2024) -
Gaining the Sparse Rewards by Exploring Lottery Tickets in Spiking Neural Network
by: Cheng, Hao, et al.
Published: (2023)