CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
Fuente:
arXiv
Saved in:
| Main Authors: | Koishigarina, Darina, Uselis, Arnas, Oh, Seong Joon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How can embedding models bind concepts?
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026)
by: Morelli, Fabian, et al.
Published: (2026)
On the rankability of visual embeddings
by: Sonthalia, Ankit, et al.
Published: (2025)
by: Sonthalia, Ankit, et al.
Published: (2025)
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026)
by: Kargi, Bora, et al.
Published: (2026)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
by: Jeong, Yujin, et al.
Published: (2025)
by: Jeong, Yujin, et al.
Published: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
Intermediate Layer Classifiers for OOD generalization
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
by: Magistri, Simone, et al.
Published: (2026)
by: Magistri, Simone, et al.
Published: (2026)
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
by: Mistretta, Marco, et al.
Published: (2025)
by: Mistretta, Marco, et al.
Published: (2025)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
by: Wang, Hu, et al.
Published: (2023)
by: Wang, Hu, et al.
Published: (2023)
Does Data Scaling Lead to Visual Compositional Generalization?
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
Multi-level Cross-modal Alignment for Image Clustering
by: Qiu, Liping, et al.
Published: (2024)
by: Qiu, Liping, et al.
Published: (2024)
Multi-modal Representation Learning for Cross-modal Prediction of Continuous Weather Patterns from Discrete Low-Dimensional Data
by: Qayyum, Alif Bin Abdul, et al.
Published: (2024)
by: Qayyum, Alif Bin Abdul, et al.
Published: (2024)
Cross-modal Active Complementary Learning with Self-refining Correspondence
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
Reconstructing facade details using MLS point clouds and Bag-of-Words approach
by: Froech, Thomas, et al.
Published: (2024)
by: Froech, Thomas, et al.
Published: (2024)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2024)
by: Wang, Enguang, et al.
Published: (2024)
Pretrained Visual Uncertainties
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
Model alignment using inter-modal bridges
by: Gholamzadeh, Ali, et al.
Published: (2025)
by: Gholamzadeh, Ali, et al.
Published: (2025)
Cross-modal feature fusion for robust point cloud registration with ambiguous geometry
by: Wang, Zhaoyi, et al.
Published: (2025)
by: Wang, Zhaoyi, et al.
Published: (2025)
Cross-modal Causal Relation Alignment for Video Question Grounding
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
by: Madaan, Divyam, et al.
Published: (2025)
by: Madaan, Divyam, et al.
Published: (2025)
Improving Text-based Person Search via Part-level Cross-modal Correspondence
by: Park, Jicheol, et al.
Published: (2024)
by: Park, Jicheol, et al.
Published: (2024)
OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models
by: Ray, Lala Shakti Swarup, et al.
Published: (2024)
by: Ray, Lala Shakti Swarup, et al.
Published: (2024)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
MEME: Multi-entity & Evolving Memory Evaluation
by: Jung, Seokwon, et al.
Published: (2026)
by: Jung, Seokwon, et al.
Published: (2026)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
by: Bai, Andrew, et al.
Published: (2025)
by: Bai, Andrew, et al.
Published: (2025)
Generalizable Single-Source Cross-modality Medical Image Segmentation via Invariant Causal Mechanisms
by: Chen, Boqi, et al.
Published: (2024)
by: Chen, Boqi, et al.
Published: (2024)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?
by: Hesse, Robin, et al.
Published: (2025)
by: Hesse, Robin, et al.
Published: (2025)
Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
by: Sun, Changchang, et al.
Published: (2025)
by: Sun, Changchang, et al.
Published: (2025)
Universal Algorithm-Implicit Learning
by: Woerner, Stefano, et al.
Published: (2026)
by: Woerner, Stefano, et al.
Published: (2026)
Multi-level and Multi-modal Action Anticipation
by: Kim, Seulgi, et al.
Published: (2025)
by: Kim, Seulgi, et al.
Published: (2025)
Towards Multi-modal Transformers in Federated Learning
by: Sun, Guangyu, et al.
Published: (2024)
by: Sun, Guangyu, et al.
Published: (2024)
Multi-modal learning for geospatial vegetation forecasting
by: Benson, Vitus, et al.
Published: (2023)
by: Benson, Vitus, et al.
Published: (2023)
Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
On the Multi-modal Vulnerability of Diffusion Models
by: Yang, Dingcheng, et al.
Published: (2024)
by: Yang, Dingcheng, et al.
Published: (2024)
Latent Object Characteristics Recognition with Visual to Haptic-Audio Cross-modal Transfer Learning
by: Saito, Namiko, et al.
Published: (2024)
by: Saito, Namiko, et al.
Published: (2024)
Multi-modal Data Binding for Survival Analysis Modeling with Incomplete Data and Annotations
by: Qu, Linhao, et al.
Published: (2024)
by: Qu, Linhao, et al.
Published: (2024)
cadrille: Multi-modal CAD Reconstruction with Reinforcement Learning
by: Kolodiazhnyi, Maksim, et al.
Published: (2025)
by: Kolodiazhnyi, Maksim, et al.
Published: (2025)
Similar Items
-
How can embedding models bind concepts?
by: Uselis, Arnas, et al.
Published: (2026) -
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026) -
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026) -
On the rankability of visual embeddings
by: Sonthalia, Ankit, et al.
Published: (2025) -
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026)