Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Uselis, Arnas, Dittadi, Andrea, Oh, Seong Joon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
por: Koishigarina, Darina, et al.
Publicado: (2025)
por: Koishigarina, Darina, et al.
Publicado: (2025)
How can embedding models bind concepts?
por: Uselis, Arnas, et al.
Publicado: (2026)
por: Uselis, Arnas, et al.
Publicado: (2026)
Does Data Scaling Lead to Visual Compositional Generalization?
por: Uselis, Arnas, et al.
Publicado: (2025)
por: Uselis, Arnas, et al.
Publicado: (2025)
Half-Truths Break Similarity-Based Retrieval
por: Kargi, Bora, et al.
Publicado: (2026)
por: Kargi, Bora, et al.
Publicado: (2026)
On the rankability of visual embeddings
por: Sonthalia, Ankit, et al.
Publicado: (2025)
por: Sonthalia, Ankit, et al.
Publicado: (2025)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
por: Jeong, Yujin, et al.
Publicado: (2025)
por: Jeong, Yujin, et al.
Publicado: (2025)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
por: Morelli, Fabian, et al.
Publicado: (2026)
por: Morelli, Fabian, et al.
Publicado: (2026)
When Do Diffusion Models learn to Generate Multiple Objects?
por: Jeong, Yujin, et al.
Publicado: (2026)
por: Jeong, Yujin, et al.
Publicado: (2026)
Intermediate Layer Classifiers for OOD generalization
por: Uselis, Arnas, et al.
Publicado: (2025)
por: Uselis, Arnas, et al.
Publicado: (2025)
Are Object-Centric Representations Better At Compositional Generalization?
por: Kapl, Ferdinand, et al.
Publicado: (2026)
por: Kapl, Ferdinand, et al.
Publicado: (2026)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
por: Mamaghan, Amir Mohammad Karimi, et al.
Publicado: (2024)
por: Mamaghan, Amir Mohammad Karimi, et al.
Publicado: (2024)
Breaking the Likelihood-Quality Trade-off in Diffusion Models by Merging Pretrained Experts
por: Esfandiari, Yasin, et al.
Publicado: (2025)
por: Esfandiari, Yasin, et al.
Publicado: (2025)
Pretrained Visual Uncertainties
por: Kirchhof, Michael, et al.
Publicado: (2024)
por: Kirchhof, Michael, et al.
Publicado: (2024)
Scalable Ensemble Diversification for OOD Generalization and Detection
por: Rubinstein, Alexander, et al.
Publicado: (2024)
por: Rubinstein, Alexander, et al.
Publicado: (2024)
DiffEnc: Variational Diffusion with a Learned Encoder
por: Nielsen, Beatrix M. G., et al.
Publicado: (2023)
por: Nielsen, Beatrix M. G., et al.
Publicado: (2023)
MEME: Multi-entity & Evolving Memory Evaluation
por: Jung, Seokwon, et al.
Publicado: (2026)
por: Jung, Seokwon, et al.
Publicado: (2026)
Universal Algorithm-Implicit Learning
por: Woerner, Stefano, et al.
Publicado: (2026)
por: Woerner, Stefano, et al.
Publicado: (2026)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
por: Berasi, Davide, et al.
Publicado: (2025)
por: Berasi, Davide, et al.
Publicado: (2025)
Are We Done with Object-Centric Learning?
por: Rubinstein, Alexander, et al.
Publicado: (2025)
por: Rubinstein, Alexander, et al.
Publicado: (2025)
SCOPE: Semantic Coreset with Orthogonal Projection Embeddings for Federated learning
por: Hossen, Md Anwar, et al.
Publicado: (2026)
por: Hossen, Md Anwar, et al.
Publicado: (2026)
Multilingual Diversity Improves Vision-Language Representations
por: Nguyen, Thao, et al.
Publicado: (2024)
por: Nguyen, Thao, et al.
Publicado: (2024)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
por: Scimeca, Luca, et al.
Publicado: (2023)
por: Scimeca, Luca, et al.
Publicado: (2023)
Assessing Neural Network Robustness via Adversarial Pivotal Tuning
por: Christensen, Peter Ebert, et al.
Publicado: (2022)
por: Christensen, Peter Ebert, et al.
Publicado: (2022)
Neighborhood-Adaptive Generalized Linear Graph Embedding with Latent Pattern Mining
por: Peng, S., et al.
Publicado: (2025)
por: Peng, S., et al.
Publicado: (2025)
Automated Learning of Semantic Embedding Representations for Diffusion Models
por: Jiang, Limai, et al.
Publicado: (2025)
por: Jiang, Limai, et al.
Publicado: (2025)
Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models
por: Kubaty, Piotr, et al.
Publicado: (2026)
por: Kubaty, Piotr, et al.
Publicado: (2026)
CORE-MTL: Rethinking Gradient Balancing via Causal Orthogonal Representations
por: Wu, Chengfeng, et al.
Publicado: (2026)
por: Wu, Chengfeng, et al.
Publicado: (2026)
Vision Foundation Model Embedding-Based Semantic Anomaly Detection
por: Ronecker, Max Peter, et al.
Publicado: (2025)
por: Ronecker, Max Peter, et al.
Publicado: (2025)
Training Without Orthogonalization, Inference With SVD: A Gradient Analysis of Rotation Representations
por: Choy, Chris
Publicado: (2026)
por: Choy, Chris
Publicado: (2026)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
por: Trager, Matthew, et al.
Publicado: (2023)
por: Trager, Matthew, et al.
Publicado: (2023)
Scratching Visual Transformer's Back with Uniform Attention
por: Hyeon-Woo, Nam, et al.
Publicado: (2022)
por: Hyeon-Woo, Nam, et al.
Publicado: (2022)
Learning Structured Representations with Hyperbolic Embeddings
por: Sinha, Aditya, et al.
Publicado: (2024)
por: Sinha, Aditya, et al.
Publicado: (2024)
Generation is Required for Data-Efficient Perception
por: Brady, Jack, et al.
Publicado: (2025)
por: Brady, Jack, et al.
Publicado: (2025)
Probing the Representational Power of Sparse Autoencoders in Vision Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
Hier-COS: Making Deep Features Hierarchy-aware via Composition of Orthogonal Subspaces
por: Sani, Depanshu, et al.
Publicado: (2025)
por: Sani, Depanshu, et al.
Publicado: (2025)
Rotary Position Embedding for Vision Transformer
por: Heo, Byeongho, et al.
Publicado: (2024)
por: Heo, Byeongho, et al.
Publicado: (2024)
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
por: Wang, Jiayun, et al.
Publicado: (2026)
por: Wang, Jiayun, et al.
Publicado: (2026)
Taming Feed-forward Reconstruction Models as Latent Encoders for 3D Generative Models
por: Wizadwongsa, Suttisak, et al.
Publicado: (2024)
por: Wizadwongsa, Suttisak, et al.
Publicado: (2024)
SITUATE: Indoor Human Trajectory Prediction through Geometric Features and Self-Supervised Vision Representation
por: Capogrosso, Luigi, et al.
Publicado: (2024)
por: Capogrosso, Luigi, et al.
Publicado: (2024)
PHyCLIP: $\ell_1$-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
por: Yoshikawa, Daiki, et al.
Publicado: (2025)
por: Yoshikawa, Daiki, et al.
Publicado: (2025)
Ejemplares similares
-
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
por: Koishigarina, Darina, et al.
Publicado: (2025) -
How can embedding models bind concepts?
por: Uselis, Arnas, et al.
Publicado: (2026) -
Does Data Scaling Lead to Visual Compositional Generalization?
por: Uselis, Arnas, et al.
Publicado: (2025) -
Half-Truths Break Similarity-Based Retrieval
por: Kargi, Bora, et al.
Publicado: (2026) -
On the rankability of visual embeddings
por: Sonthalia, Ankit, et al.
Publicado: (2025)