Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Balakrishnan, Ravikumar, Mendapara, Sanket, Garg, Ankit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations
por: Balakrishnan, Ravikumar, et al.
Publicado: (2026)
por: Balakrishnan, Ravikumar, et al.
Publicado: (2026)
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
por: Waseda, Futa, et al.
Publicado: (2025)
por: Waseda, Futa, et al.
Publicado: (2025)
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
por: Ying, Zonghao, et al.
Publicado: (2026)
por: Ying, Zonghao, et al.
Publicado: (2026)
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
por: Balakrishnan, Ravikumar, et al.
Publicado: (2025)
por: Balakrishnan, Ravikumar, et al.
Publicado: (2025)
VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
por: Phute, Mansi, et al.
Publicado: (2025)
por: Phute, Mansi, et al.
Publicado: (2025)
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
por: Cheng, Hao, et al.
Publicado: (2024)
por: Cheng, Hao, et al.
Publicado: (2024)
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
por: Zhu, Jiayi, et al.
Publicado: (2025)
por: Zhu, Jiayi, et al.
Publicado: (2025)
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
por: Cheng, Hao, et al.
Publicado: (2024)
por: Cheng, Hao, et al.
Publicado: (2024)
Typographic Text Generation with Off-the-Shelf Diffusion Model
por: Peong, KhayTze, et al.
Publicado: (2024)
por: Peong, KhayTze, et al.
Publicado: (2024)
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
por: Jose, Cijo, et al.
Publicado: (2024)
por: Jose, Cijo, et al.
Publicado: (2024)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
por: Li, Yueyan, et al.
Publicado: (2025)
por: Li, Yueyan, et al.
Publicado: (2025)
Automatic Text Box Placement for Supporting Typographic Design
por: Muraoka, Jun, et al.
Publicado: (2025)
por: Muraoka, Jun, et al.
Publicado: (2025)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
por: Qraitem, Maan, et al.
Publicado: (2024)
por: Qraitem, Maan, et al.
Publicado: (2024)
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
por: Johnson, Emily, et al.
Publicado: (2025)
por: Johnson, Emily, et al.
Publicado: (2025)
In the Era of Prompt Learning with Vision-Language Models
por: Jha, Ankit
Publicado: (2024)
por: Jha, Ankit
Publicado: (2024)
Fit Pixels, Get Labels: Meta-learned Implicit Networks for Image Segmentation
por: Vyas, Kushal, et al.
Publicado: (2025)
por: Vyas, Kushal, et al.
Publicado: (2025)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
por: He, Zoe Wanying, et al.
Publicado: (2025)
por: He, Zoe Wanying, et al.
Publicado: (2025)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
por: Hufe, Lorenz, et al.
Publicado: (2025)
por: Hufe, Lorenz, et al.
Publicado: (2025)
SineProject: Machine Unlearning for Stable Vision Language Alignment
por: Garg, Arpit, et al.
Publicado: (2025)
por: Garg, Arpit, et al.
Publicado: (2025)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
por: Chen, Tianle, et al.
Publicado: (2026)
por: Chen, Tianle, et al.
Publicado: (2026)
SGHA-Attack: Semantic-Guided Hierarchical Alignment for Transferable Targeted Attacks on Vision-Language Models
por: Wang, Haobo, et al.
Publicado: (2026)
por: Wang, Haobo, et al.
Publicado: (2026)
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
por: Sharma, Pranav, et al.
Publicado: (2025)
por: Sharma, Pranav, et al.
Publicado: (2025)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
por: Cao, Yue, et al.
Publicado: (2024)
por: Cao, Yue, et al.
Publicado: (2024)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
por: Singh, Ishika, et al.
Publicado: (2025)
por: Singh, Ishika, et al.
Publicado: (2025)
Revisiting Vision Language Foundations for No-Reference Image Quality Assessment
por: Yadav, Ankit, et al.
Publicado: (2025)
por: Yadav, Ankit, et al.
Publicado: (2025)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
Language-Image Alignment with Fixed Text Encoders
por: Yang, Jingfeng, et al.
Publicado: (2025)
por: Yang, Jingfeng, et al.
Publicado: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
por: Liang, Wenqi, et al.
Publicado: (2025)
por: Liang, Wenqi, et al.
Publicado: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
por: Xu, Jingning, et al.
Publicado: (2026)
por: Xu, Jingning, et al.
Publicado: (2026)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
por: Liu, Yang, et al.
Publicado: (2025)
por: Liu, Yang, et al.
Publicado: (2025)
Image Recognition with Vision and Language Embeddings of VLMs
por: Volkov, Illia, et al.
Publicado: (2025)
por: Volkov, Illia, et al.
Publicado: (2025)
Pixel Is Not a Barrier: An Effective Evasion Attack for Pixel-Domain Diffusion Models
por: Shih, Chun-Yen, et al.
Publicado: (2024)
por: Shih, Chun-Yen, et al.
Publicado: (2024)
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation
por: Li, Bingyu, et al.
Publicado: (2025)
por: Li, Bingyu, et al.
Publicado: (2025)
Detecting Text Manipulation in Images using Vision Language Models
por: Vidit, Vidit, et al.
Publicado: (2025)
por: Vidit, Vidit, et al.
Publicado: (2025)
Embedding Textual Information in Images Using Quinary Pixel Combinations
por: Kandala, A V Uday Kiran
Publicado: (2026)
por: Kandala, A V Uday Kiran
Publicado: (2026)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
por: Liu, Zhiheng, et al.
Publicado: (2026)
por: Liu, Zhiheng, et al.
Publicado: (2026)
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
por: Ma, Xianzhi, et al.
Publicado: (2025)
por: Ma, Xianzhi, et al.
Publicado: (2025)
SPOOF: Simple Pixel Operations for Out-of-Distribution Fooling
por: Gupta, Ankit, et al.
Publicado: (2025)
por: Gupta, Ankit, et al.
Publicado: (2025)
Reading Between the Lanes: Text VideoQA on the Road
por: Tom, George, et al.
Publicado: (2023)
por: Tom, George, et al.
Publicado: (2023)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
por: Bao, Muyi, et al.
Publicado: (2026)
por: Bao, Muyi, et al.
Publicado: (2026)
Ejemplares similares
-
One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations
por: Balakrishnan, Ravikumar, et al.
Publicado: (2026) -
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
por: Waseda, Futa, et al.
Publicado: (2025) -
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
por: Ying, Zonghao, et al.
Publicado: (2026) -
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
por: Balakrishnan, Ravikumar, et al.
Publicado: (2025) -
VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
por: Phute, Mansi, et al.
Publicado: (2025)