Extract Free Dense Misalignment from CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Nam, JeongYeon, Im, Jinbae, Kim, Wonjae, Kil, Taeho |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EGTR: Extracting Graph from Transformer for Scene Graph Generation
by: Im, Jinbae, et al.
Published: (2024)
by: Im, Jinbae, et al.
Published: (2024)
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
by: Paik, Gio, et al.
Published: (2025)
by: Paik, Gio, et al.
Published: (2025)
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024)
by: Chun, Sanghyuk, et al.
Published: (2024)
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
by: Mistretta, Marco, et al.
Published: (2025)
by: Mistretta, Marco, et al.
Published: (2025)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
Towards Reliable Test-Time Adaptation: Style Invariance as a Correctness Likelihood
by: Nam, Gilhyun, et al.
Published: (2025)
by: Nam, Gilhyun, et al.
Published: (2025)
One-Cycle Structured Pruning via Stability-Driven Subnetwork Search
by: Ghimire, Deepak, et al.
Published: (2025)
by: Ghimire, Deepak, et al.
Published: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
by: Lee, Sua, et al.
Published: (2026)
by: Lee, Sua, et al.
Published: (2026)
ConDL: Detector-Free Dense Image Matching
by: Kwiatkowski, Monika, et al.
Published: (2024)
by: Kwiatkowski, Monika, et al.
Published: (2024)
DeCLIP: Decoding CLIP representations for deepfake localization
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
kNN-CLIP: Retrieval Enables Training-Free Segmentation on Continually Expanding Large Vocabularies
by: Gui, Zhongrui, et al.
Published: (2024)
by: Gui, Zhongrui, et al.
Published: (2024)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
by: Magistri, Simone, et al.
Published: (2026)
by: Magistri, Simone, et al.
Published: (2026)
Enhancing CLIP with CLIP: Exploring Pseudolabeling for Limited-Label Prompt Tuning
by: Menghini, Cristina, et al.
Published: (2023)
by: Menghini, Cristina, et al.
Published: (2023)
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)
by: Kim, Sohee, et al.
Published: (2024)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
by: Cai, Yichao, et al.
Published: (2025)
by: Cai, Yichao, et al.
Published: (2025)
FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
by: Wei, Xiyuan, et al.
Published: (2025)
by: Wei, Xiyuan, et al.
Published: (2025)
Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning
by: Hümmer, Christoph, et al.
Published: (2023)
by: Hümmer, Christoph, et al.
Published: (2023)
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
by: Yang, Zhiwei, et al.
Published: (2025)
by: Yang, Zhiwei, et al.
Published: (2025)
FastCLIP: A Suite of Optimization Techniques to Accelerate CLIP Training with Limited Resources
by: Wei, Xiyuan, et al.
Published: (2024)
by: Wei, Xiyuan, et al.
Published: (2024)
Breaking the Limits of Open-Weight CLIP: An Optimization Framework for Self-supervised Fine-tuning of CLIP
by: Mehta, Anant, et al.
Published: (2026)
by: Mehta, Anant, et al.
Published: (2026)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)
by: Wang, Xinze, et al.
Published: (2025)
MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental Learning
by: Nicolas, Julien, et al.
Published: (2023)
by: Nicolas, Julien, et al.
Published: (2023)
TiC-CLIP: Continual Training of CLIP Models
by: Garg, Saurabh, et al.
Published: (2023)
by: Garg, Saurabh, et al.
Published: (2023)
Online Zero-Shot Classification with CLIP
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
MIP: CLIP-based Image Reconstruction from PEFT Gradients
by: Zhou, Peiheng, et al.
Published: (2024)
by: Zhou, Peiheng, et al.
Published: (2024)
What do we learn from inverting CLIP models?
by: Kazemi, Hamid, et al.
Published: (2024)
by: Kazemi, Hamid, et al.
Published: (2024)
Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
by: Kim, Bumjun, et al.
Published: (2026)
by: Kim, Bumjun, et al.
Published: (2026)
X2CT-CLIP: Enable Multi-Abnormality Detection in Computed Tomography from Chest Radiography via Tri-Modal Contrastive Learning
by: You, Jianzhong, et al.
Published: (2025)
by: You, Jianzhong, et al.
Published: (2025)
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
by: Lee, Junwon, et al.
Published: (2024)
by: Lee, Junwon, et al.
Published: (2024)
Finetuning CLIP to Reason about Pairwise Differences
by: Sam, Dylan, et al.
Published: (2024)
by: Sam, Dylan, et al.
Published: (2024)
Detecting AI-Generated Images via CLIP
by: Moskowitz, A. G., et al.
Published: (2024)
by: Moskowitz, A. G., et al.
Published: (2024)
Is CLIP ideal? No. Can we fix it? Yes!
by: Kang, Raphi, et al.
Published: (2025)
by: Kang, Raphi, et al.
Published: (2025)
HyperCLIP: Adapting Vision-Language models with Hypernetworks
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective
by: Zhang, Yanan, et al.
Published: (2024)
by: Zhang, Yanan, et al.
Published: (2024)
Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning
by: Peleg, Amit, et al.
Published: (2025)
by: Peleg, Amit, et al.
Published: (2025)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
by: Ahmad, Shahzad, et al.
Published: (2023)
by: Ahmad, Shahzad, et al.
Published: (2023)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Noise Projection: Closing the Prompt-Agnostic Gap Behind Text-to-Image Misalignment in Diffusion Models
by: Tong, Yunze, et al.
Published: (2025)
by: Tong, Yunze, et al.
Published: (2025)
Similar Items
-
EGTR: Extracting Graph from Transformer for Scene Graph Generation
by: Im, Jinbae, et al.
Published: (2024) -
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
by: Paik, Gio, et al.
Published: (2025) -
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024) -
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
by: Mistretta, Marco, et al.
Published: (2025) -
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)