CLIP's Visual Embedding Projector is a Few-shot Cornucopia
Fuente:
arXiv
Saved in:
| Main Authors: | Fahes, Mohammad, Vu, Tuan-Hung, Bursuc, Andrei, Pérez, Patrick, de Charette, Raoul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Domain Adaptation with a Single Vision-Language Embedding
by: Fahes, Mohammad, et al.
Published: (2024)
by: Fahes, Mohammad, et al.
Published: (2024)
A Simple Recipe for Language-guided Domain Generalized Segmentation
by: Fahes, Mohammad, et al.
Published: (2023)
by: Fahes, Mohammad, et al.
Published: (2023)
FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2025)
by: Benigmim, Yasser, et al.
Published: (2025)
DenseMTL: Cross-task Attention Mechanism for Dense Multi-task Learning
by: Lopes, Ivan, et al.
Published: (2022)
by: Lopes, Ivan, et al.
Published: (2022)
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
by: Wysoczańska, Monika, et al.
Published: (2023)
by: Wysoczańska, Monika, et al.
Published: (2023)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024)
by: Cao, Anh-Quan, et al.
Published: (2024)
MadCLIP: Few-shot Medical Anomaly Detection with CLIP
by: Shiri, Mahshid, et al.
Published: (2025)
by: Shiri, Mahshid, et al.
Published: (2025)
MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection
by: Zhang, Ximiao, et al.
Published: (2024)
by: Zhang, Ximiao, et al.
Published: (2024)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
by: Jon, Hyo Jin, et al.
Published: (2026)
by: Jon, Hyo Jin, et al.
Published: (2026)
IQE-CLIP: Instance-aware Query Embedding for Zero-/Few-shot Anomaly Detection in Medical Domain
by: Huang, Hong, et al.
Published: (2025)
by: Huang, Hong, et al.
Published: (2025)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
by: Du, Yao, et al.
Published: (2025)
by: Du, Yao, et al.
Published: (2025)
Material Transforms from Disentangled NeRF Representations
by: Lopes, Ivan, et al.
Published: (2024)
by: Lopes, Ivan, et al.
Published: (2024)
DiffCLIP: Few-shot Language-driven Multimodal Classifier
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
CLIP-guided Prototype Modulating for Few-shot Action Recognition
by: Wang, Xiang, et al.
Published: (2023)
by: Wang, Xiang, et al.
Published: (2023)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
by: Magistri, Simone, et al.
Published: (2026)
by: Magistri, Simone, et al.
Published: (2026)
PaSCo: Urban 3D Panoptic Scene Completion with Uncertainty Awareness
by: Cao, Anh-Quan, et al.
Published: (2023)
by: Cao, Anh-Quan, et al.
Published: (2023)
Selective Vision-Language Subspace Projection for Few-shot CLIP
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
CRoF: CLIP-based Robust Few-shot Learning on Noisy Labels
by: Deng, Shizhuo, et al.
Published: (2024)
by: Deng, Shizhuo, et al.
Published: (2024)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
Contrastive Bi-Projector for Unsupervised Domain Adaption
by: Huang, Lin-Chieh, et al.
Published: (2023)
by: Huang, Lin-Chieh, et al.
Published: (2023)
MatSwap: Light-aware material transfers in images
by: Lopes, Ivan, et al.
Published: (2025)
by: Lopes, Ivan, et al.
Published: (2025)
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
by: Sirko-Galouchenko, Sophia, et al.
Published: (2025)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2025)
DREAM: Visual Decoding from Reversing Human Visual System
by: Xia, Weihao, et al.
Published: (2023)
by: Xia, Weihao, et al.
Published: (2023)
OccAny: Generalized Unconstrained Urban 3D Occupancy
by: Cao, Anh-Quan, et al.
Published: (2026)
by: Cao, Anh-Quan, et al.
Published: (2026)
StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets
by: Cao, Anh-Quan, et al.
Published: (2025)
by: Cao, Anh-Quan, et al.
Published: (2025)
LiDPM: Rethinking Point Diffusion for Lidar Scene Completion
by: Martyniuk, Tetiana, et al.
Published: (2025)
by: Martyniuk, Tetiana, et al.
Published: (2025)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024)
by: Vobecky, Antonin, et al.
Published: (2024)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
by: Vobecky, Antonin, et al.
Published: (2022)
by: Vobecky, Antonin, et al.
Published: (2022)
The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning
by: Li, Yaohui, et al.
Published: (2024)
by: Li, Yaohui, et al.
Published: (2024)
BIGFix: Bidirectional Image Generation with Token Fixing
by: Besnier, Victor, et al.
Published: (2025)
by: Besnier, Victor, et al.
Published: (2025)
FLIER: Few-shot Language Image Models Embedded with Latent Representations
by: Zhou, Zhinuo, et al.
Published: (2024)
by: Zhou, Zhinuo, et al.
Published: (2024)
UMBRAE: Unified Multimodal Brain Decoding
by: Xia, Weihao, et al.
Published: (2024)
by: Xia, Weihao, et al.
Published: (2024)
TokenPacker: Efficient Visual Projector for Multimodal LLM
by: Li, Wentong, et al.
Published: (2024)
by: Li, Wentong, et al.
Published: (2024)
Reliability in Semantic Segmentation: Can We Use Synthetic Data?
by: Loiseau, Thibaut, et al.
Published: (2023)
by: Loiseau, Thibaut, et al.
Published: (2023)
Retrieval-Enhanced Visual Prompt Learning for Few-shot Classification
by: Rong, Jintao, et al.
Published: (2023)
by: Rong, Jintao, et al.
Published: (2023)
Multi-Perspective Data Augmentation for Few-shot Object Detection
by: Vu, Anh-Khoa Nguyen, et al.
Published: (2025)
by: Vu, Anh-Khoa Nguyen, et al.
Published: (2025)
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
by: Kim, Donghyeong, et al.
Published: (2025)
by: Kim, Donghyeong, et al.
Published: (2025)
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
by: Tang, Yuwei, et al.
Published: (2024)
by: Tang, Yuwei, et al.
Published: (2024)
Breaking the Scalability Limit of Multi-Projector Calibration with Embedded Cameras
by: Kawano, Takumi, et al.
Published: (2026)
by: Kawano, Takumi, et al.
Published: (2026)
Three Pillars improving Vision Foundation Model Distillation for Lidar
by: Puy, Gilles, et al.
Published: (2023)
by: Puy, Gilles, et al.
Published: (2023)
Similar Items
-
Domain Adaptation with a Single Vision-Language Embedding
by: Fahes, Mohammad, et al.
Published: (2024) -
A Simple Recipe for Language-guided Domain Generalized Segmentation
by: Fahes, Mohammad, et al.
Published: (2023) -
FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2025) -
DenseMTL: Cross-task Attention Mechanism for Dense Multi-task Learning
by: Lopes, Ivan, et al.
Published: (2022) -
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
by: Wysoczańska, Monika, et al.
Published: (2023)