Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Lu, Zhang, Haiyang, Xu, Changsheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Complementary Text-Guided Attention for Zero-Shot Adversarial Robustness
di: Yu, Lu, et al.
Pubblicazione: (2026)
di: Yu, Lu, et al.
Pubblicazione: (2026)
Text is All You Need for Vision-Language Model Jailbreaking
di: Chen, Yihang, et al.
Pubblicazione: (2026)
di: Chen, Yihang, et al.
Pubblicazione: (2026)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
di: Zhao, Kairan, et al.
Pubblicazione: (2026)
di: Zhao, Kairan, et al.
Pubblicazione: (2026)
FineVision: Open Data Is All You Need
di: Wiedmann, Luis, et al.
Pubblicazione: (2025)
di: Wiedmann, Luis, et al.
Pubblicazione: (2025)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
di: Lin, Feng, et al.
Pubblicazione: (2025)
di: Lin, Feng, et al.
Pubblicazione: (2025)
Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models
di: Xu, Zane, et al.
Pubblicazione: (2025)
di: Xu, Zane, et al.
Pubblicazione: (2025)
LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment
di: Hu, Junyi, et al.
Pubblicazione: (2026)
di: Hu, Junyi, et al.
Pubblicazione: (2026)
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
di: Xu, Changhua, et al.
Pubblicazione: (2026)
di: Xu, Changhua, et al.
Pubblicazione: (2026)
[MASK] is All You Need
di: Hu, Vincent Tao, et al.
Pubblicazione: (2024)
di: Hu, Vincent Tao, et al.
Pubblicazione: (2024)
Ideal Registration? Segmentation is All You Need
di: Chen, Xiang, et al.
Pubblicazione: (2025)
di: Chen, Xiang, et al.
Pubblicazione: (2025)
Zoom and Shift are All You Need
di: Qin, Jiahao
Pubblicazione: (2024)
di: Qin, Jiahao
Pubblicazione: (2024)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
di: Zhang, Pu, et al.
Pubblicazione: (2025)
di: Zhang, Pu, et al.
Pubblicazione: (2025)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
Memory augment is All You Need for image restoration
di: Zhang, Xiao Feng, et al.
Pubblicazione: (2023)
di: Zhang, Xiao Feng, et al.
Pubblicazione: (2023)
Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language Models
di: Chen, Mengyuan, et al.
Pubblicazione: (2024)
di: Chen, Mengyuan, et al.
Pubblicazione: (2024)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
di: Jeong, Seongjun, et al.
Pubblicazione: (2024)
di: Jeong, Seongjun, et al.
Pubblicazione: (2024)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
di: Bendou, Yassir, et al.
Pubblicazione: (2024)
di: Bendou, Yassir, et al.
Pubblicazione: (2024)
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
di: Wang, Haibo, et al.
Pubblicazione: (2026)
di: Wang, Haibo, et al.
Pubblicazione: (2026)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
di: Cao, Pu, et al.
Pubblicazione: (2023)
di: Cao, Pu, et al.
Pubblicazione: (2023)
AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models
di: Cui, Yubo, et al.
Pubblicazione: (2026)
di: Cui, Yubo, et al.
Pubblicazione: (2026)
Binary Verification for Zero-Shot Vision
di: Hu, Rongbin, et al.
Pubblicazione: (2025)
di: Hu, Rongbin, et al.
Pubblicazione: (2025)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
di: Yang, Cheng, et al.
Pubblicazione: (2026)
di: Yang, Cheng, et al.
Pubblicazione: (2026)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
di: Zhang, Naifu, et al.
Pubblicazione: (2025)
di: Zhang, Naifu, et al.
Pubblicazione: (2025)
Fast Wrong-way Cycling Detection in CCTV Videos: Sparse Sampling is All You Need
di: Xu, Jing, et al.
Pubblicazione: (2024)
di: Xu, Jing, et al.
Pubblicazione: (2024)
Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
di: Zhang, Letian, et al.
Pubblicazione: (2025)
di: Zhang, Letian, et al.
Pubblicazione: (2025)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
di: Xu, Sirui, et al.
Pubblicazione: (2024)
di: Xu, Sirui, et al.
Pubblicazione: (2024)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
di: Ye, Wei, et al.
Pubblicazione: (2024)
di: Ye, Wei, et al.
Pubblicazione: (2024)
Language-Driven Anchors for Zero-Shot Adversarial Robustness
di: Li, Xiao, et al.
Pubblicazione: (2023)
di: Li, Xiao, et al.
Pubblicazione: (2023)
Attention Prompting on Image for Large Vision-Language Models
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets
di: Choi, Lucas, et al.
Pubblicazione: (2024)
di: Choi, Lucas, et al.
Pubblicazione: (2024)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
di: Sajib, Rakib Hossain, et al.
Pubblicazione: (2026)
di: Sajib, Rakib Hossain, et al.
Pubblicazione: (2026)
Xihe: Scalable Zero-Shot Time Series Learner Via Hierarchical Interleaved Block Attention
di: Sun, Yinbo, et al.
Pubblicazione: (2025)
di: Sun, Yinbo, et al.
Pubblicazione: (2025)
Camouflaged Image Synthesis Is All You Need to Boost Camouflaged Detection
di: Zhang, Haichao, et al.
Pubblicazione: (2023)
di: Zhang, Haichao, et al.
Pubblicazione: (2023)
TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability
di: Ma, Fengji, et al.
Pubblicazione: (2024)
di: Ma, Fengji, et al.
Pubblicazione: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
di: Luo, Kun, et al.
Pubblicazione: (2026)
di: Luo, Kun, et al.
Pubblicazione: (2026)
All You Need in Knowledge Distillation Is a Tailored Coordinate System
di: Zhou, Junjie, et al.
Pubblicazione: (2024)
di: Zhou, Junjie, et al.
Pubblicazione: (2024)
Rethinking Deep Clustering Paradigms: Self-Supervision Is All You Need
di: Shaheena, Amal, et al.
Pubblicazione: (2025)
di: Shaheena, Amal, et al.
Pubblicazione: (2025)
TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation
di: Li, Dingbang, et al.
Pubblicazione: (2024)
di: Li, Dingbang, et al.
Pubblicazione: (2024)
Off-The-Shelf Image-to-Image Models Are All You Need To Defeat Image Protection Schemes
di: Pleimling, Xavier, et al.
Pubblicazione: (2026)
di: Pleimling, Xavier, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Complementary Text-Guided Attention for Zero-Shot Adversarial Robustness
di: Yu, Lu, et al.
Pubblicazione: (2026) -
Text is All You Need for Vision-Language Model Jailbreaking
di: Chen, Yihang, et al.
Pubblicazione: (2026) -
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
di: Zhao, Kairan, et al.
Pubblicazione: (2026) -
FineVision: Open Data Is All You Need
di: Wiedmann, Luis, et al.
Pubblicazione: (2025) -
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
di: Lin, Feng, et al.
Pubblicazione: (2025)