How can embedding models bind concepts?
Fuente:
arXiv
Saved in:
| Main Authors: | Uselis, Arnas, Koishigarina, Darina, Oh, Seong Joon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
by: Koishigarina, Darina, et al.
Published: (2025)
by: Koishigarina, Darina, et al.
Published: (2025)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
On the rankability of visual embeddings
by: Sonthalia, Ankit, et al.
Published: (2025)
by: Sonthalia, Ankit, et al.
Published: (2025)
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026)
by: Kargi, Bora, et al.
Published: (2026)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026)
by: Morelli, Fabian, et al.
Published: (2026)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
by: Jeong, Yujin, et al.
Published: (2025)
by: Jeong, Yujin, et al.
Published: (2025)
Intermediate Layer Classifiers for OOD generalization
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
Does Data Scaling Lead to Visual Compositional Generalization?
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
Pretrained Visual Uncertainties
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
MEME: Multi-entity & Evolving Memory Evaluation
by: Jung, Seokwon, et al.
Published: (2026)
by: Jung, Seokwon, et al.
Published: (2026)
Universal Algorithm-Implicit Learning
by: Woerner, Stefano, et al.
Published: (2026)
by: Woerner, Stefano, et al.
Published: (2026)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
by: Scimeca, Luca, et al.
Published: (2023)
by: Scimeca, Luca, et al.
Published: (2023)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution
by: Park, Yonghyun, et al.
Published: (2025)
by: Park, Yonghyun, et al.
Published: (2025)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
by: Janizek, Joseph D., et al.
Published: (2026)
by: Janizek, Joseph D., et al.
Published: (2026)
Fine-tuning can cripple your foundation model; preserving features may be the solution
by: Mukhoti, Jishnu, et al.
Published: (2023)
by: Mukhoti, Jishnu, et al.
Published: (2023)
AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data
by: Brown, Christopher F., et al.
Published: (2025)
by: Brown, Christopher F., et al.
Published: (2025)
Bayesian generative models can flag performance loss, bias, and out-of-distribution image content
by: López-Pérez, Miguel, et al.
Published: (2025)
by: López-Pérez, Miguel, et al.
Published: (2025)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
Sparse autoencoders reveal selective remapping of visual concepts during adaptation
by: Lim, Hyesu, et al.
Published: (2024)
by: Lim, Hyesu, et al.
Published: (2024)
A Kolmogorov metric embedding for live cell microscopy signaling patterns
by: Aho, Layton, et al.
Published: (2024)
by: Aho, Layton, et al.
Published: (2024)
WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local Views
by: Jin, SeongMin, et al.
Published: (2026)
by: Jin, SeongMin, et al.
Published: (2026)
Attention Frequency Modulation: Training-Free Spectral Modulation of Diffusion Cross-Attention
by: Oh, Seunghun, et al.
Published: (2026)
by: Oh, Seunghun, et al.
Published: (2026)
Beyond Softmax: Dual-Branch Sigmoid Architecture for Accurate Class Activation Maps
by: Oh, Yoojin, et al.
Published: (2025)
by: Oh, Yoojin, et al.
Published: (2025)
Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
by: Shen, Shufan, et al.
Published: (2025)
by: Shen, Shufan, et al.
Published: (2025)
Wide Two-Layer Networks can Learn from Adversarial Perturbations
by: Kumano, Soichiro, et al.
Published: (2024)
by: Kumano, Soichiro, et al.
Published: (2024)
Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
Taming Feed-forward Reconstruction Models as Latent Encoders for 3D Generative Models
by: Wizadwongsa, Suttisak, et al.
Published: (2024)
by: Wizadwongsa, Suttisak, et al.
Published: (2024)
How to build a consistency model: Learning flow maps via self-distillation
by: Boffi, Nicholas M., et al.
Published: (2025)
by: Boffi, Nicholas M., et al.
Published: (2025)
Analyzing Effects of Mixed Sample Data Augmentation on Model Interpretability
by: Won, Soyoun, et al.
Published: (2023)
by: Won, Soyoun, et al.
Published: (2023)
Zeros can be Informative: Masked Binary U-Net for Image Segmentation on Tensor Cores
by: Wu, Chunshu, et al.
Published: (2026)
by: Wu, Chunshu, et al.
Published: (2026)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
Towards scientific discovery with dictionary learning: Extracting biological concepts from microscopy foundation models
by: Donhauser, Konstantin, et al.
Published: (2024)
by: Donhauser, Konstantin, et al.
Published: (2024)
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026)
by: Seong, Hyun Seok, et al.
Published: (2026)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026)
by: Moon, WonJun, et al.
Published: (2026)
Parallel Tempering Initial Sampling in Inference-Time Reward Alignment
by: Oh, Myeongjun, et al.
Published: (2026)
by: Oh, Myeongjun, et al.
Published: (2026)
Generative Spatiotemporal Data Augmentation
by: Zhou, Jinfan, et al.
Published: (2025)
by: Zhou, Jinfan, et al.
Published: (2025)
Similar Items
-
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
by: Koishigarina, Darina, et al.
Published: (2025) -
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026) -
On the rankability of visual embeddings
by: Sonthalia, Ankit, et al.
Published: (2025) -
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026) -
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026)