WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ning, Shan, Qiu, Longtian, Sun, Jiaxuan, He, Xuming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
von: Ning, Shan, et al.
Veröffentlicht: (2026)
von: Ning, Shan, et al.
Veröffentlicht: (2026)
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
von: Qiu, Longtian, et al.
Veröffentlicht: (2025)
von: Qiu, Longtian, et al.
Veröffentlicht: (2025)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
von: Qiu, Longtian, et al.
Veröffentlicht: (2024)
von: Qiu, Longtian, et al.
Veröffentlicht: (2024)
Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning
von: Zhou, Hongkuan, et al.
Veröffentlicht: (2025)
von: Zhou, Hongkuan, et al.
Veröffentlicht: (2025)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2026)
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2026)
Part-Aware Open-Vocabulary 3D Affordance Grounding via Prototypical Semantic and Geometric Alignment
von: Gou, Dongqiang, et al.
Veröffentlicht: (2026)
von: Gou, Dongqiang, et al.
Veröffentlicht: (2026)
A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
von: Dai, Ying, et al.
Veröffentlicht: (2025)
von: Dai, Ying, et al.
Veröffentlicht: (2025)
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
von: Alyami, Sarah, et al.
Veröffentlicht: (2025)
von: Alyami, Sarah, et al.
Veröffentlicht: (2025)
Exploring Open-Vocabulary Object Recognition in Images using CLIP
von: Chen, Wei Yu, et al.
Veröffentlicht: (2026)
von: Chen, Wei Yu, et al.
Veröffentlicht: (2026)
Grounding Language Models for Visual Entity Recognition
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
von: Song, Dan, et al.
Veröffentlicht: (2023)
von: Song, Dan, et al.
Veröffentlicht: (2023)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
von: He, Chiyuan, et al.
Veröffentlicht: (2025)
von: He, Chiyuan, et al.
Veröffentlicht: (2025)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection
von: Zhang, Ximiao, et al.
Veröffentlicht: (2024)
von: Zhang, Ximiao, et al.
Veröffentlicht: (2024)
Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC
von: Hu, Guanyu, et al.
Veröffentlicht: (2025)
von: Hu, Guanyu, et al.
Veröffentlicht: (2025)
Freeze and Cluster: A Simple Baseline for Rehearsal-Free Continual Category Discovery
von: Zhang, Chuyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chuyu, et al.
Veröffentlicht: (2025)
Cross-domain EEG-based Emotion Recognition with Contrastive Learning
von: Yan, Rui, et al.
Veröffentlicht: (2025)
von: Yan, Rui, et al.
Veröffentlicht: (2025)
Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher Learning
von: Li, Yan, et al.
Veröffentlicht: (2023)
von: Li, Yan, et al.
Veröffentlicht: (2023)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
von: Yan, Yibin, et al.
Veröffentlicht: (2024)
von: Yan, Yibin, et al.
Veröffentlicht: (2024)
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
von: Meng, GuangHao, et al.
Veröffentlicht: (2025)
von: Meng, GuangHao, et al.
Veröffentlicht: (2025)
UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting
von: Kim, Geonuk, et al.
Veröffentlicht: (2026)
von: Kim, Geonuk, et al.
Veröffentlicht: (2026)
Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP
von: Nie, Sen, et al.
Veröffentlicht: (2026)
von: Nie, Sen, et al.
Veröffentlicht: (2026)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
von: Zhou, Jinxin, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxin, et al.
Veröffentlicht: (2025)
Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
big.LITTLE Vision Transformer for Efficient Visual Recognition
von: Guo, He, et al.
Veröffentlicht: (2024)
von: Guo, He, et al.
Veröffentlicht: (2024)
Emotion Recognition with CLIP and Sequential Learning
von: Zhou, Weiwei, et al.
Veröffentlicht: (2025)
von: Zhou, Weiwei, et al.
Veröffentlicht: (2025)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
von: Zhu, Wenqi, et al.
Veröffentlicht: (2024)
von: Zhu, Wenqi, et al.
Veröffentlicht: (2024)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
von: Ning, Shan, et al.
Veröffentlicht: (2026) -
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
von: Qiu, Longtian, et al.
Veröffentlicht: (2025) -
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
von: Qiu, Longtian, et al.
Veröffentlicht: (2024) -
Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning
von: Zhou, Hongkuan, et al.
Veröffentlicht: (2025) -
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)