Saved in:
| Main Authors: | Hu, Junyi, Zhou, Qiji, Zhang, Lei, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.08156 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
by: Zhou, Qiji, et al.
Published: (2024)
by: Zhou, Qiji, et al.
Published: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
by: Zhou, Yuxi, et al.
Published: (2026)
by: Zhou, Yuxi, et al.
Published: (2026)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
by: Jiang, Yankai, et al.
Published: (2024)
by: Jiang, Yankai, et al.
Published: (2024)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
by: Zhang, Erhang, et al.
Published: (2025)
by: Zhang, Erhang, et al.
Published: (2025)
Language-Driven Visual Consensus for Zero-Shot Semantic Segmentation
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
by: Qiu, Longtian, et al.
Published: (2024)
by: Qiu, Longtian, et al.
Published: (2024)
AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
by: Qu, Zhen, et al.
Published: (2026)
by: Qu, Zhen, et al.
Published: (2026)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
Language-Driven Anchors for Zero-Shot Adversarial Robustness
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
by: Xu, Sirui, et al.
Published: (2024)
by: Xu, Sirui, et al.
Published: (2024)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
by: Xu, Wenjia, et al.
Published: (2024)
by: Xu, Wenjia, et al.
Published: (2024)
Ontology-Guided Diffusion for Zero-Shot Visual Sim2Real Transfer
by: Youssef, Mohamed, et al.
Published: (2026)
by: Youssef, Mohamed, et al.
Published: (2026)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
by: Gao, Haihan, et al.
Published: (2024)
by: Gao, Haihan, et al.
Published: (2024)
Pseudo-label Based Domain Adaptation for Zero-Shot Text Steganalysis
by: Luo, Yufei, et al.
Published: (2024)
by: Luo, Yufei, et al.
Published: (2024)
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
by: Zhang, Jielu, et al.
Published: (2023)
by: Zhang, Jielu, et al.
Published: (2023)
VGTS: Visually Guided Text Spotting for Novel Categories in Historical Manuscripts
by: Hu, Wenbo, et al.
Published: (2023)
by: Hu, Wenbo, et al.
Published: (2023)
Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
by: Zhang, Qian, et al.
Published: (2025)
by: Zhang, Qian, et al.
Published: (2025)
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
by: Byun, Ji Young, et al.
Published: (2025)
by: Byun, Ji Young, et al.
Published: (2025)
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
by: Yang, Ziteng, et al.
Published: (2025)
by: Yang, Ziteng, et al.
Published: (2025)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
by: Wu, Yuanli, et al.
Published: (2025)
by: Wu, Yuanli, et al.
Published: (2025)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)
by: Lin, Jiawen, et al.
Published: (2025)
Coherent Zero-Shot Visual Instruction Generation
by: Phung, Quynh, et al.
Published: (2024)
by: Phung, Quynh, et al.
Published: (2024)
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
by: Shivika, et al.
Published: (2026)
by: Shivika, et al.
Published: (2026)
Visual Language Models as Zero-Shot Deepfake Detectors
by: Pirogov, Viacheslav
Published: (2025)
by: Pirogov, Viacheslav
Published: (2025)
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models
by: Cui, Yubo, et al.
Published: (2026)
by: Cui, Yubo, et al.
Published: (2026)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Binary Verification for Zero-Shot Vision
by: Hu, Rongbin, et al.
Published: (2025)
by: Hu, Rongbin, et al.
Published: (2025)
Leveraging Unknown Objects to Construct Labeled-Unlabeled Meta-Relationships for Zero-Shot Object Navigation
by: Zheng, Yanwei, et al.
Published: (2024)
by: Zheng, Yanwei, et al.
Published: (2024)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
by: Zhao, Lirui, et al.
Published: (2024)
by: Zhao, Lirui, et al.
Published: (2024)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
by: Liu, Jeffrey, et al.
Published: (2025)
by: Liu, Jeffrey, et al.
Published: (2025)
Language-Guided Structure-Aware Network for Camouflaged Object Detection
by: Zhang, Min
Published: (2026)
by: Zhang, Min
Published: (2026)
AdaNeg: Adaptive Negative Proxy Guided OOD Detection with Vision-Language Models
by: Zhang, Yabin, et al.
Published: (2024)
by: Zhang, Yabin, et al.
Published: (2024)
TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
by: Wilson, Bibin
Published: (2026)
by: Wilson, Bibin
Published: (2026)
RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
by: Li, Ke, et al.
Published: (2025)
by: Li, Ke, et al.
Published: (2025)
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
by: Pan, Li, et al.
Published: (2024)
by: Pan, Li, et al.
Published: (2024)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
by: Zhang, Pu, et al.
Published: (2025)
by: Zhang, Pu, et al.
Published: (2025)
Similar Items
-
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
by: Zhou, Qiji, et al.
Published: (2024) -
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
by: Yu, Lu, et al.
Published: (2024) -
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
by: Zhou, Yuxi, et al.
Published: (2026) -
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
by: Jiang, Yankai, et al.
Published: (2024) -
Zero-Shot Temporal Interaction Localization for Egocentric Videos
by: Zhang, Erhang, et al.
Published: (2025)