LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Anh-Quan, Jaritz, Maximilian, Guillaumin, Matthieu, de Charette, Raoul, Bazzani, Loris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniCoRN: Unified Commented Retrieval Network with LMMs
von: Jaritz, Maximilian, et al.
Veröffentlicht: (2025)
von: Jaritz, Maximilian, et al.
Veröffentlicht: (2025)
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2025)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
CLIP's Visual Embedding Projector is a Few-shot Cornucopia
von: Fahes, Mohammad, et al.
Veröffentlicht: (2024)
von: Fahes, Mohammad, et al.
Veröffentlicht: (2024)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
von: Asokan, Mothilal, et al.
Veröffentlicht: (2025)
von: Asokan, Mothilal, et al.
Veröffentlicht: (2025)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
PaSCo: Urban 3D Panoptic Scene Completion with Uncertainty Awareness
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2023)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2023)
LMM-Regularized CLIP Embeddings for Image Classification
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
Demystifying CLIP Data
von: Xu, Hu, et al.
Veröffentlicht: (2023)
von: Xu, Hu, et al.
Veröffentlicht: (2023)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
von: Balauca, Ada-Astrid, et al.
Veröffentlicht: (2024)
von: Balauca, Ada-Astrid, et al.
Veröffentlicht: (2024)
TiC-CLIP: Continual Training of CLIP Models
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
von: Silva, Sathira, et al.
Veröffentlicht: (2025)
von: Silva, Sathira, et al.
Veröffentlicht: (2025)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
Jina CLIP: Your CLIP Model Is Also Your Text Retriever
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
TLAC: Two-stage LMM Augmented CLIP for Zero-Shot Classification
von: Munir, Ans, et al.
Veröffentlicht: (2025)
von: Munir, Ans, et al.
Veröffentlicht: (2025)
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
von: Xia, Rui, et al.
Veröffentlicht: (2025)
von: Xia, Rui, et al.
Veröffentlicht: (2025)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
von: Tang, Yuwei, et al.
Veröffentlicht: (2024)
von: Tang, Yuwei, et al.
Veröffentlicht: (2024)
UMBRAE: Unified Multimodal Brain Decoding
von: Xia, Weihao, et al.
Veröffentlicht: (2024)
von: Xia, Weihao, et al.
Veröffentlicht: (2024)
CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values
von: Koleilat, Taha, et al.
Veröffentlicht: (2025)
von: Koleilat, Taha, et al.
Veröffentlicht: (2025)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Meta CLIP 2: A Worldwide Scaling Recipe
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
Generalizable Prompt Learning of CLIP: A Brief Overview
von: Cui, Fangming, et al.
Veröffentlicht: (2025)
von: Cui, Fangming, et al.
Veröffentlicht: (2025)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
UniCoRN: Unified Commented Retrieval Network with LMMs
von: Jaritz, Maximilian, et al.
Veröffentlicht: (2025) -
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025) -
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024) -
StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2025) -
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)