Hierarchically Robust Zero-shot Vision-language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Junhao, Zhang, Yifei, Zhu, Hao, Ong, Yew-Soon, Koniusz, Piotr |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Possibilistic Predictive Uncertainty for Deep Learning
by: Ni, Yao, et al.
Published: (2026)
by: Ni, Yao, et al.
Published: (2026)
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition
by: Lin, Shen, et al.
Published: (2026)
by: Lin, Shen, et al.
Published: (2026)
Feature Hallucination for Self-supervised Action Recognition
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
SiamNAS: Siamese Surrogate Model for Dominance Relation Prediction in Multi-objective Neural Architecture Search
by: Zhou, Yuyang, et al.
Published: (2025)
by: Zhou, Yuyang, et al.
Published: (2025)
Enhancing Adversarial Robustness via Uncertainty-Aware Distributional Adversarial Training
by: Dong, Junhao, et al.
Published: (2024)
by: Dong, Junhao, et al.
Published: (2024)
Video Understanding by Design: How Datasets Shape Architectures and Insights
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Pre-training with Random Orthogonal Projection Image Modeling
by: Haghighat, Maryam, et al.
Published: (2023)
by: Haghighat, Maryam, et al.
Published: (2023)
Uncertainty-DTW for Sequences and Visual Tokens
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
by: Xie, Jiahao, et al.
Published: (2023)
by: Xie, Jiahao, et al.
Published: (2023)
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
by: Ding, Dexuan, et al.
Published: (2024)
by: Ding, Dexuan, et al.
Published: (2024)
MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
Subspace Kernel Learning on Tensor Sequences
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
Graph Your Own Prompt
by: Ding, Xi, et al.
Published: (2025)
by: Ding, Xi, et al.
Published: (2025)
Learning Time in Static Classifiers
by: Ding, Xi, et al.
Published: (2025)
by: Ding, Xi, et al.
Published: (2025)
Adaptive Multi-head Contrastive Learning
by: Wang, Lei, et al.
Published: (2023)
by: Wang, Lei, et al.
Published: (2023)
Generative AI-based Prompt Evolution Engineering Design Optimization With Vision-Language Model
by: Wong, Melvin, et al.
Published: (2024)
by: Wong, Melvin, et al.
Published: (2024)
Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal-Viewpoint Alignment
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Zero-shot Concept Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
by: Yang, Chenglin, et al.
Published: (2023)
by: Yang, Chenglin, et al.
Published: (2023)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
by: Liu, Junming, et al.
Published: (2026)
by: Liu, Junming, et al.
Published: (2026)
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
by: Kim, Ye-Chan, et al.
Published: (2025)
by: Kim, Ye-Chan, et al.
Published: (2025)
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
by: Ni, Yao, et al.
Published: (2025)
by: Ni, Yao, et al.
Published: (2025)
LLM2TEA: An Agentic AI Designer for Discovery with Generative Evolutionary Multitasking
by: Wong, Melvin, et al.
Published: (2024)
by: Wong, Melvin, et al.
Published: (2024)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
by: Zhao, Heng, et al.
Published: (2026)
by: Zhao, Heng, et al.
Published: (2026)
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
by: Stanić, Aleksandar, et al.
Published: (2024)
by: Stanić, Aleksandar, et al.
Published: (2024)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
by: Jeong, Hyeonho, et al.
Published: (2023)
by: Jeong, Hyeonho, et al.
Published: (2023)
CHAIN: Enhancing Generalization in Data-Efficient GANs via lipsCHitz continuity constrAIned Normalization
by: Ni, Yao, et al.
Published: (2024)
by: Ni, Yao, et al.
Published: (2024)
Prompt Evolution for Generative AI: A Classifier-Guided Approach
by: Wong, Melvin, et al.
Published: (2023)
by: Wong, Melvin, et al.
Published: (2023)
SimSAM: Zero-shot Medical Image Segmentation via Simulated Interaction
by: Towle, Benjamin, et al.
Published: (2024)
by: Towle, Benjamin, et al.
Published: (2024)
Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation
by: Mistretta, Marco, et al.
Published: (2024)
by: Mistretta, Marco, et al.
Published: (2024)
Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly Detection
by: Zhu, Jiawen, et al.
Published: (2024)
by: Zhu, Jiawen, et al.
Published: (2024)
To Trust Or Not To Trust Your Vision-Language Model's Prediction
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Language-Driven Anchors for Zero-Shot Adversarial Robustness
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
by: Schlarmann, Christian, et al.
Published: (2024)
by: Schlarmann, Christian, et al.
Published: (2024)
Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
by: Hong, Yunqi, et al.
Published: (2025)
by: Hong, Yunqi, et al.
Published: (2025)
VaPR -- Vision-language Preference alignment for Reasoning
by: Wadhawan, Rohan, et al.
Published: (2025)
by: Wadhawan, Rohan, et al.
Published: (2025)
Similar Items
-
Possibilistic Predictive Uncertainty for Deep Learning
by: Ni, Yao, et al.
Published: (2026) -
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition
by: Lin, Shen, et al.
Published: (2026) -
Feature Hallucination for Self-supervised Action Recognition
by: Wang, Lei, et al.
Published: (2025) -
SiamNAS: Siamese Surrogate Model for Dominance Relation Prediction in Multi-objective Neural Architecture Search
by: Zhou, Yuyang, et al.
Published: (2025) -
Enhancing Adversarial Robustness via Uncertainty-Aware Distributional Adversarial Training
by: Dong, Junhao, et al.
Published: (2024)