Is CLIP ideal? No. Can we fix it? Yes!
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Raphi, Song, Yue, Gkioxari, Georgia, Perona, Pietro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
by: Kang, Raphi, et al.
Published: (2026)
by: Kang, Raphi, et al.
Published: (2026)
Unsupervised Representation Learning from Sparse Transformation Analysis
by: Song, Yue, et al.
Published: (2024)
by: Song, Yue, et al.
Published: (2024)
Single View Seafloor Recovery from Imaging Sonar via Differentiable Rendering
by: Brodjian, Sevan, et al.
Published: (2026)
by: Brodjian, Sevan, et al.
Published: (2026)
Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations
by: Fogliato, Riccardo, et al.
Published: (2023)
by: Fogliato, Riccardo, et al.
Published: (2023)
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)
by: Kim, Sohee, et al.
Published: (2024)
Kuramoto Orientation Diffusion Models
by: Song, Yue, et al.
Published: (2025)
by: Song, Yue, et al.
Published: (2025)
What do we learn from inverting CLIP models?
by: Kazemi, Hamid, et al.
Published: (2024)
by: Kazemi, Hamid, et al.
Published: (2024)
Find Any Part in 3D
by: Ma, Ziqi, et al.
Published: (2024)
by: Ma, Ziqi, et al.
Published: (2024)
Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
by: Sahoo, Aadarsh, et al.
Published: (2026)
by: Sahoo, Aadarsh, et al.
Published: (2026)
Representational Difference Explanations
by: Kondapaneni, Neehar, et al.
Published: (2025)
by: Kondapaneni, Neehar, et al.
Published: (2025)
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Adversarially Robust CLIP Models Can Induce Better (Robust) Perceptual Metrics
by: Croce, Francesco, et al.
Published: (2025)
by: Croce, Francesco, et al.
Published: (2025)
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025)
by: Sahoo, Aadarsh, et al.
Published: (2025)
Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
by: Demler, Ilona, et al.
Published: (2025)
by: Demler, Ilona, et al.
Published: (2025)
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
by: Ma, Ziqi, et al.
Published: (2026)
by: Ma, Ziqi, et al.
Published: (2026)
Social Perception of Faces in a Vision-Language Model
by: Hausladen, Carina I., et al.
Published: (2024)
by: Hausladen, Carina I., et al.
Published: (2024)
DeCLIP: Decoding CLIP representations for deepfake localization
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
by: Magistri, Simone, et al.
Published: (2026)
by: Magistri, Simone, et al.
Published: (2026)
Enhancing CLIP with CLIP: Exploring Pseudolabeling for Limited-Label Prompt Tuning
by: Menghini, Cristina, et al.
Published: (2023)
by: Menghini, Cristina, et al.
Published: (2023)
NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
by: Wei, Xiyuan, et al.
Published: (2025)
by: Wei, Xiyuan, et al.
Published: (2025)
FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
FastCLIP: A Suite of Optimization Techniques to Accelerate CLIP Training with Limited Resources
by: Wei, Xiyuan, et al.
Published: (2024)
by: Wei, Xiyuan, et al.
Published: (2024)
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025)
by: Ma, Ziqi, et al.
Published: (2025)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)
by: Wang, Xinze, et al.
Published: (2025)
Breaking the Limits of Open-Weight CLIP: An Optimization Framework for Self-supervised Fine-tuning of CLIP
by: Mehta, Anant, et al.
Published: (2026)
by: Mehta, Anant, et al.
Published: (2026)
MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental Learning
by: Nicolas, Julien, et al.
Published: (2023)
by: Nicolas, Julien, et al.
Published: (2023)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
TiC-CLIP: Continual Training of CLIP Models
by: Garg, Saurabh, et al.
Published: (2023)
by: Garg, Saurabh, et al.
Published: (2023)
CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning
by: Frascaroli, Emanuele, et al.
Published: (2024)
by: Frascaroli, Emanuele, et al.
Published: (2024)
Online Zero-Shot Classification with CLIP
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
Finetuning CLIP to Reason about Pairwise Differences
by: Sam, Dylan, et al.
Published: (2024)
by: Sam, Dylan, et al.
Published: (2024)
Extract Free Dense Misalignment from CLIP
by: Nam, JeongYeon, et al.
Published: (2024)
by: Nam, JeongYeon, et al.
Published: (2024)
Detecting AI-Generated Images via CLIP
by: Moskowitz, A. G., et al.
Published: (2024)
by: Moskowitz, A. G., et al.
Published: (2024)
Reconstructing Hand-Held Objects in 3D from Images and Videos
by: Wu, Jane, et al.
Published: (2024)
by: Wu, Jane, et al.
Published: (2024)
Align and Distill: Unifying and Improving Domain Adaptive Object Detection
by: Kay, Justin, et al.
Published: (2024)
by: Kay, Justin, et al.
Published: (2024)
Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning
by: Peleg, Amit, et al.
Published: (2025)
by: Peleg, Amit, et al.
Published: (2025)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
HyperCLIP: Adapting Vision-Language models with Hypernetworks
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
by: Ahmad, Shahzad, et al.
Published: (2023)
by: Ahmad, Shahzad, et al.
Published: (2023)
Similar Items
-
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
by: Kang, Raphi, et al.
Published: (2026) -
Unsupervised Representation Learning from Sparse Transformation Analysis
by: Song, Yue, et al.
Published: (2024) -
Single View Seafloor Recovery from Imaging Sonar via Differentiable Rendering
by: Brodjian, Sevan, et al.
Published: (2026) -
Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations
by: Fogliato, Riccardo, et al.
Published: (2023) -
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)