FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhuang, Jiedong, Hu, Jiaqi, Mu, Lianrui, Hu, Rui, Liang, Xiaoyu, Ye, Jiangnan, Hu, Haoji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
von: Liang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liang, Xiaoyu, et al.
Veröffentlicht: (2024)
TeG-DG: Textually Guided Domain Generalization for Face Anti-Spoofing
von: Mu, Lianrui, et al.
Veröffentlicht: (2023)
von: Mu, Lianrui, et al.
Veröffentlicht: (2023)
Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting
von: Ye, Jiangnan, et al.
Veröffentlicht: (2025)
von: Ye, Jiangnan, et al.
Veröffentlicht: (2025)
No Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image Detection
von: Mu, Lianrui, et al.
Veröffentlicht: (2025)
von: Mu, Lianrui, et al.
Veröffentlicht: (2025)
Identity-Preserving Pose-Guided Character Animation via Facial Landmarks Transformation
von: Mu, Lianrui, et al.
Veröffentlicht: (2024)
von: Mu, Lianrui, et al.
Veröffentlicht: (2024)
Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2026)
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2026)
ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
Online Zero-Shot Classification with CLIP
von: Qian, Qi, et al.
Veröffentlicht: (2024)
von: Qian, Qi, et al.
Veröffentlicht: (2024)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
CLAIR: CLIP-Aided Weakly Supervised Zero-Shot Cross-Domain Image Retrieval
von: Tan, Chor Boon, et al.
Veröffentlicht: (2025)
von: Tan, Chor Boon, et al.
Veröffentlicht: (2025)
Making Better Mistakes in CLIP-Based Zero-Shot Classification with Hierarchy-Aware Language Prompts
von: Liang, Tong, et al.
Veröffentlicht: (2025)
von: Liang, Tong, et al.
Veröffentlicht: (2025)
AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection
von: Cao, Yunkang, et al.
Veröffentlicht: (2024)
von: Cao, Yunkang, et al.
Veröffentlicht: (2024)
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
von: Hu, Ming, et al.
Veröffentlicht: (2026)
von: Hu, Ming, et al.
Veröffentlicht: (2026)
Human Scanpath Prediction in Target-Present Visual Search with Semantic-Foveal Bayesian Attention
von: Luzio, João, et al.
Veröffentlicht: (2025)
von: Luzio, João, et al.
Veröffentlicht: (2025)
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
von: Yin, Xiaojie, et al.
Veröffentlicht: (2025)
von: Yin, Xiaojie, et al.
Veröffentlicht: (2025)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
Enhancing Zero-Shot Anomaly Detection: CLIP-SAM Collaboration with Cascaded Prompts
von: Hou, Yanning, et al.
Veröffentlicht: (2025)
von: Hou, Yanning, et al.
Veröffentlicht: (2025)
PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
von: Chen, Zewen, et al.
Veröffentlicht: (2024)
von: Chen, Zewen, et al.
Veröffentlicht: (2024)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
Robustness-Guided Image Synthesis for Data-Free Quantization
von: Bai, Jianhong, et al.
Veröffentlicht: (2023)
von: Bai, Jianhong, et al.
Veröffentlicht: (2023)
CLIP-driven Zero-shot Learning with Ambiguous Labels
von: Fan, Jinfu, et al.
Veröffentlicht: (2026)
von: Fan, Jinfu, et al.
Veröffentlicht: (2026)
StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection
von: Hou, Yanning, et al.
Veröffentlicht: (2025)
von: Hou, Yanning, et al.
Veröffentlicht: (2025)
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
von: Yang, Ziteng, et al.
Veröffentlicht: (2025)
von: Yang, Ziteng, et al.
Veröffentlicht: (2025)
Towards Distribution-Agnostic Generalized Category Discovery
von: Bai, Jianhong, et al.
Veröffentlicht: (2023)
von: Bai, Jianhong, et al.
Veröffentlicht: (2023)
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
von: Bai, Jianhong, et al.
Veröffentlicht: (2025)
von: Bai, Jianhong, et al.
Veröffentlicht: (2025)
Transductive Zero-Shot and Few-Shot CLIP
von: Martin, Ségolène, et al.
Veröffentlicht: (2024)
von: Martin, Ségolène, et al.
Veröffentlicht: (2024)
DeCLIP: Decoupled Prompting for CLIP-based Multi-Label Class-Incremental Learning
von: Du, Kaile, et al.
Veröffentlicht: (2025)
von: Du, Kaile, et al.
Veröffentlicht: (2025)
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
von: Kim, Donghyeong, et al.
Veröffentlicht: (2025)
von: Kim, Donghyeong, et al.
Veröffentlicht: (2025)
Dynamic Token Reduction during Generation for Vision Language Models
von: Liang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Liang, Xiaoyu, et al.
Veröffentlicht: (2025)
RobuSTereo: Robust Zero-Shot Stereo Matching under Adverse Weather
von: Wang, Yuran, et al.
Veröffentlicht: (2025)
von: Wang, Yuran, et al.
Veröffentlicht: (2025)
WMoE-CLIP: Wavelet-Enhanced Mixture-of-Experts Prompt Learning for Zero-Shot Anomaly Detection
von: Chen, Peng, et al.
Veröffentlicht: (2026)
von: Chen, Peng, et al.
Veröffentlicht: (2026)
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
von: Liu, Man, et al.
Veröffentlicht: (2024)
von: Liu, Man, et al.
Veröffentlicht: (2024)
Binary Verification for Zero-Shot Vision
von: Hu, Rongbin, et al.
Veröffentlicht: (2025)
von: Hu, Rongbin, et al.
Veröffentlicht: (2025)
Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC
von: Hu, Guanyu, et al.
Veröffentlicht: (2025)
von: Hu, Guanyu, et al.
Veröffentlicht: (2025)
Zero-Shot Class Unlearning in CLIP with Synthetic Samples
von: Kravets, A., et al.
Veröffentlicht: (2024)
von: Kravets, A., et al.
Veröffentlicht: (2024)
AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation
von: Fang, Qingqing, et al.
Veröffentlicht: (2025)
von: Fang, Qingqing, et al.
Veröffentlicht: (2025)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
von: Ali, Muhammad, et al.
Veröffentlicht: (2024)
von: Ali, Muhammad, et al.
Veröffentlicht: (2024)
StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
von: Hu, Jinghao, et al.
Veröffentlicht: (2026)
von: Hu, Jinghao, et al.
Veröffentlicht: (2026)
Visual Adaptive Prompting for Compositional Zero-Shot Learning
von: Stein, Kyle, et al.
Veröffentlicht: (2025)
von: Stein, Kyle, et al.
Veröffentlicht: (2025)
Dual-Image Enhanced CLIP for Zero-Shot Anomaly Detection
von: Zhang, Zhaoxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaoxiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
von: Liang, Xiaoyu, et al.
Veröffentlicht: (2024) -
TeG-DG: Textually Guided Domain Generalization for Face Anti-Spoofing
von: Mu, Lianrui, et al.
Veröffentlicht: (2023) -
Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting
von: Ye, Jiangnan, et al.
Veröffentlicht: (2025) -
No Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image Detection
von: Mu, Lianrui, et al.
Veröffentlicht: (2025) -
Identity-Preserving Pose-Guided Character Animation via Facial Landmarks Transformation
von: Mu, Lianrui, et al.
Veröffentlicht: (2024)