Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Hui, Wang, Wenya, Chen, Kecheng, Liu, Jie, Liu, Yibing, Qin, Tiexin, He, Peisong, Jiang, Xinghao, Li, Haoliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognition
by: Liu, Hui, et al.
Published: (2026)
by: Liu, Hui, et al.
Published: (2026)
Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start Scenarios
by: Liu, Hui, et al.
Published: (2026)
by: Liu, Hui, et al.
Published: (2026)
Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need
by: Chen, Kecheng, et al.
Published: (2024)
by: Chen, Kecheng, et al.
Published: (2024)
Test-time Adaptation for Foundation Medical Segmentation Model without Parametric Updates
by: Chen, Kecheng, et al.
Published: (2025)
by: Chen, Kecheng, et al.
Published: (2025)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
by: Liu, Hui, et al.
Published: (2023)
by: Liu, Hui, et al.
Published: (2023)
Test-time adaptation for image compression with distribution regularization
by: Chen, Kecheng, et al.
Published: (2024)
by: Chen, Kecheng, et al.
Published: (2024)
Generalizing to New Dynamical Systems via Frequency Domain Adaptation
by: Qin, Tiexin, et al.
Published: (2025)
by: Qin, Tiexin, et al.
Published: (2025)
TELLER: A Trustworthy Framework for Explainable, Generalizable and Controllable Fake News Detection
by: Liu, Hui, et al.
Published: (2024)
by: Liu, Hui, et al.
Published: (2024)
Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction Regression
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
by: Zhang, Keyang, et al.
Published: (2025)
by: Zhang, Keyang, et al.
Published: (2025)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models
by: Chen, Kecheng, et al.
Published: (2026)
by: Chen, Kecheng, et al.
Published: (2026)
Learning Dynamic Graph Embeddings with Neural Controlled Differential Equations
by: Qin, Tiexin, et al.
Published: (2023)
by: Qin, Tiexin, et al.
Published: (2023)
Zero-Shot Visual Concept Blending Without Text Guidance
by: Makino, Hiroya, et al.
Published: (2025)
by: Makino, Hiroya, et al.
Published: (2025)
CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
by: Shou, Zhou-Peng, et al.
Published: (2025)
by: Shou, Zhou-Peng, et al.
Published: (2025)
Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
by: Liu, Ting, et al.
Published: (2025)
by: Liu, Ting, et al.
Published: (2025)
Efficient Test-Time Adaptation through Latent Subspace Coefficients Search
by: Luo, Xinyu, et al.
Published: (2025)
by: Luo, Xinyu, et al.
Published: (2025)
EZSR: Event-based Zero-Shot Recognition
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
Binary Verification for Zero-Shot Vision
by: Hu, Rongbin, et al.
Published: (2025)
by: Hu, Rongbin, et al.
Published: (2025)
MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
by: Wang, Yingying, et al.
Published: (2025)
by: Wang, Yingying, et al.
Published: (2025)
VETime: Vision Enhanced Zero-Shot Time Series Anomaly Detection
by: Yang, Yingyuan, et al.
Published: (2026)
by: Yang, Yingyuan, et al.
Published: (2026)
Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models
by: Xiong, Lexiang, et al.
Published: (2025)
by: Xiong, Lexiang, et al.
Published: (2025)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
by: Lin, Hehai, et al.
Published: (2025)
by: Lin, Hehai, et al.
Published: (2025)
Visuo-Tactile Zero-Shot Object Recognition with Vision-Language Model
by: Ueda, Shiori, et al.
Published: (2024)
by: Ueda, Shiori, et al.
Published: (2024)
Deep Signature: Characterization of Large-Scale Molecular Dynamics
by: Qin, Tiexin, et al.
Published: (2024)
by: Qin, Tiexin, et al.
Published: (2024)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
by: Lu, Zimao, et al.
Published: (2025)
by: Lu, Zimao, et al.
Published: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
by: Chen, Kecheng, et al.
Published: (2025)
by: Chen, Kecheng, et al.
Published: (2025)
Hybrid Fusion: One-Minute Efficient Training for Zero-Shot Cross-Domain Image Fusion
by: Zhang, Ran, et al.
Published: (2026)
by: Zhang, Ran, et al.
Published: (2026)
Exposing AI-generated Videos: A Benchmark Dataset and a Local-and-Global Temporal Defect Based Detection Method
by: He, Peisong, et al.
Published: (2024)
by: He, Peisong, et al.
Published: (2024)
Zero-Shot CFC: Fast Real-World Image Denoising based on Cross-Frequency Consistency
by: Jiang, Yanlin, et al.
Published: (2025)
by: Jiang, Yanlin, et al.
Published: (2025)
Learning from Observer Gaze:Zero-Shot Attention Prediction Oriented by Human-Object Interaction Recognition
by: Zhou, Yuchen, et al.
Published: (2024)
by: Zhou, Yuchen, et al.
Published: (2024)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
by: Xu, Zhenlin, et al.
Published: (2023)
by: Xu, Zhenlin, et al.
Published: (2023)
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
by: He, Xuanhua, et al.
Published: (2024)
by: He, Xuanhua, et al.
Published: (2024)
Self-Improving for Zero-Shot Named Entity Recognition with Large Language Models
by: Xie, Tingyu, et al.
Published: (2023)
by: Xie, Tingyu, et al.
Published: (2023)
ZePo: Zero-Shot Portrait Stylization with Faster Sampling
by: Liu, Jin, et al.
Published: (2024)
by: Liu, Jin, et al.
Published: (2024)
CodeUnlearn: Amortized Zero-Shot Machine Unlearning in Language Models Using Discrete Concept
by: Wu, YuXuan, et al.
Published: (2024)
by: Wu, YuXuan, et al.
Published: (2024)
SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
by: Song, Jiajun, et al.
Published: (2025)
by: Song, Jiajun, et al.
Published: (2025)
PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment
by: Liu, Zhendong, et al.
Published: (2024)
by: Liu, Zhendong, et al.
Published: (2024)
Similar Items
-
Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognition
by: Liu, Hui, et al.
Published: (2026) -
Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start Scenarios
by: Liu, Hui, et al.
Published: (2026) -
Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need
by: Chen, Kecheng, et al.
Published: (2024) -
Test-time Adaptation for Foundation Medical Segmentation Model without Parametric Updates
by: Chen, Kecheng, et al.
Published: (2025) -
Interpretable Multimodal Misinformation Detection with Logic Reasoning
by: Liu, Hui, et al.
Published: (2023)