Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Waseda, Futa, Sugawara, Saku, Echizen, Isao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
by: Waseda, Futa, et al.
Published: (2024)
by: Waseda, Futa, et al.
Published: (2024)
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
by: Waseda, Futa, et al.
Published: (2025)
by: Waseda, Futa, et al.
Published: (2025)
Uncolorable Examples: Preventing Unauthorized AI Colorization via Perception-Aware Chroma-Restrictive Perturbation
by: Nii, Yuki, et al.
Published: (2025)
by: Nii, Yuki, et al.
Published: (2025)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
by: Yamabe, Shojiro, et al.
Published: (2025)
by: Yamabe, Shojiro, et al.
Published: (2025)
Defending Against Physical Adversarial Patch Attacks on Infrared Human Detection
by: Strack, Lukas, et al.
Published: (2023)
by: Strack, Lukas, et al.
Published: (2023)
Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks
by: Chengyu, Jia, et al.
Published: (2026)
by: Chengyu, Jia, et al.
Published: (2026)
Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
Exploring Self-Supervised Vision Transformers for Deepfake Detection: A Comparative Analysis
by: Nguyen, Huy H., et al.
Published: (2024)
by: Nguyen, Huy H., et al.
Published: (2024)
Do Vision-Language Foundational models show Robust Visual Perception?
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
by: Li, Jinlong, et al.
Published: (2024)
by: Li, Jinlong, et al.
Published: (2024)
Analysing the Robustness of Vision-Language-Models to Common Corruptions
by: Usama, Muhammad, et al.
Published: (2025)
by: Usama, Muhammad, et al.
Published: (2025)
On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
Anchor-based Robust Finetuning of Vision-Language Models
by: Han, Jinwei, et al.
Published: (2024)
by: Han, Jinwei, et al.
Published: (2024)
Proxy Robustness in Vision Language Models is Effortlessly Transferable
by: Fu, Xiaowei, et al.
Published: (2026)
by: Fu, Xiaowei, et al.
Published: (2026)
Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning
by: Huang, Zhuo, et al.
Published: (2023)
by: Huang, Zhuo, et al.
Published: (2023)
Same or Not? Enhancing Visual Perception in Vision-Language Models
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Rethinking Invariance Regularization in Adversarial Training to Improve Robustness-Accuracy Trade-off
by: Waseda, Futa, et al.
Published: (2024)
by: Waseda, Futa, et al.
Published: (2024)
Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation
by: Hoche, Joseph, et al.
Published: (2026)
by: Hoche, Joseph, et al.
Published: (2026)
Robust Calibration of Large Vision-Language Adapters
by: Murugesan, Balamurali, et al.
Published: (2024)
by: Murugesan, Balamurali, et al.
Published: (2024)
Dynamic Token Reweighting for Robust Vision-Language Models
by: Jiang, Tanqiu, et al.
Published: (2025)
by: Jiang, Tanqiu, et al.
Published: (2025)
What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models
by: Nie, Sen, et al.
Published: (2026)
by: Nie, Sen, et al.
Published: (2026)
Evaluating Robustness of Vision-Language Models Under Noisy Conditions
by: Purushoth, et al.
Published: (2025)
by: Purushoth, et al.
Published: (2025)
Harnessing Large Language and Vision-Language Models for Robust Out-of-Distribution Detection
by: Lee, Pei-Kang, et al.
Published: (2025)
by: Lee, Pei-Kang, et al.
Published: (2025)
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026)
by: Liu, Chao, et al.
Published: (2026)
Fine-Tuning Text-To-Image Diffusion Models for Class-Wise Spurious Feature Generation
by: MaungMaung, AprilPyone, et al.
Published: (2024)
by: MaungMaung, AprilPyone, et al.
Published: (2024)
FOCoOp: Enhancing Out-of-Distribution Robustness in Federated Prompt Learning for Vision-Language Models
by: Liao, Xinting, et al.
Published: (2025)
by: Liao, Xinting, et al.
Published: (2025)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
GFT-GCN: Privacy-Preserving 3D Face Mesh Recognition with Spectral Diffusion
by: Felouat, Hichem, et al.
Published: (2025)
by: Felouat, Hichem, et al.
Published: (2025)
Trajectory-Diversity-Driven Robust Vision-and-Language Navigation
by: Li, Jiangyang, et al.
Published: (2026)
by: Li, Jiangyang, et al.
Published: (2026)
Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise
by: Gao, Yansheng, et al.
Published: (2025)
by: Gao, Yansheng, et al.
Published: (2025)
Adversarial Robustness Analysis of Vision-Language Models in Medical Image Segmentation
by: Budathoki, Anjila, et al.
Published: (2025)
by: Budathoki, Anjila, et al.
Published: (2025)
Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models
by: Liu, Xiao, et al.
Published: (2026)
by: Liu, Xiao, et al.
Published: (2026)
HeBA: Heterogeneous Bottleneck Adapters for Robust Vision-Language Models
by: Islam, Md Jahidul
Published: (2026)
by: Islam, Md Jahidul
Published: (2026)
AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models
by: Li, Zhiwei, et al.
Published: (2026)
by: Li, Zhiwei, et al.
Published: (2026)
Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective
by: Zhou, Wanqi, et al.
Published: (2024)
by: Zhou, Wanqi, et al.
Published: (2024)
Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation
by: Ji, Yuheng, et al.
Published: (2024)
by: Ji, Yuheng, et al.
Published: (2024)
Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models
by: Nakayama, Aya, et al.
Published: (2025)
by: Nakayama, Aya, et al.
Published: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
by: Xu, Jingning, et al.
Published: (2026)
by: Xu, Jingning, et al.
Published: (2026)
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
by: Saxena, Rohit, et al.
Published: (2026)
by: Saxena, Rohit, et al.
Published: (2026)
Similar Items
-
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
by: Waseda, Futa, et al.
Published: (2024) -
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
by: Waseda, Futa, et al.
Published: (2025) -
Uncolorable Examples: Preventing Unauthorized AI Colorization via Perception-Aware Chroma-Restrictive Perturbation
by: Nii, Yuki, et al.
Published: (2025) -
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
by: Yamabe, Shojiro, et al.
Published: (2025) -
Defending Against Physical Adversarial Patch Attacks on Infrared Human Detection
by: Strack, Lukas, et al.
Published: (2023)