On the rankability of visual embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Sonthalia, Ankit, Uselis, Arnas, Oh, Seong Joon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026)
by: Morelli, Fabian, et al.
Published: (2026)
How can embedding models bind concepts?
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026)
by: Kargi, Bora, et al.
Published: (2026)
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
by: Koishigarina, Darina, et al.
Published: (2025)
by: Koishigarina, Darina, et al.
Published: (2025)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
by: Jeong, Yujin, et al.
Published: (2025)
by: Jeong, Yujin, et al.
Published: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
Intermediate Layer Classifiers for OOD generalization
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
by: Kim, Nayeong, et al.
Published: (2025)
by: Kim, Nayeong, et al.
Published: (2025)
Does Data Scaling Lead to Visual Compositional Generalization?
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO
by: Chun, Sanghyuk, et al.
Published: (2022)
by: Chun, Sanghyuk, et al.
Published: (2022)
OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation
by: Hwang, Dongjun, et al.
Published: (2024)
by: Hwang, Dongjun, et al.
Published: (2024)
Matricial Free Energy as a Gaussianizing Regularizer: Enhancing Autoencoders for Gaussian Code Generation
by: Sonthalia, Rishi, et al.
Published: (2025)
by: Sonthalia, Rishi, et al.
Published: (2025)
Pretrained Visual Uncertainties
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
Perception Encoder: The best visual embeddings are not at the output of the network
by: Bolya, Daniel, et al.
Published: (2025)
by: Bolya, Daniel, et al.
Published: (2025)
Enhancing Multi-Image Understanding through Delimiter Token Scaling
by: Lee, Minyoung, et al.
Published: (2026)
by: Lee, Minyoung, et al.
Published: (2026)
Universal Algorithm-Implicit Learning
by: Woerner, Stefano, et al.
Published: (2026)
by: Woerner, Stefano, et al.
Published: (2026)
Explaining generative diffusion models via visual analysis for interpretable decision-making process
by: Park, Ji-Hoon, et al.
Published: (2024)
by: Park, Ji-Hoon, et al.
Published: (2024)
TestDG: Test-time Domain Generalization for Continual Test-time Adaptation
by: Lee, Sohyun, et al.
Published: (2025)
by: Lee, Sohyun, et al.
Published: (2025)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
MEME: Multi-entity & Evolving Memory Evaluation
by: Jung, Seokwon, et al.
Published: (2026)
by: Jung, Seokwon, et al.
Published: (2026)
Explorer: Robust Collection of Interactable GUI Elements
by: Chaimalas, Iason, et al.
Published: (2025)
by: Chaimalas, Iason, et al.
Published: (2025)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
by: Scimeca, Luca, et al.
Published: (2023)
by: Scimeca, Luca, et al.
Published: (2023)
Elevating Flow-Guided Video Inpainting with Reference Generation
by: Cho, Suhwan, et al.
Published: (2024)
by: Cho, Suhwan, et al.
Published: (2024)
Limitations of NERF with pre-trained Vision Features for Few-Shot 3D Reconstruction
by: Sanjyal, Ankit
Published: (2025)
by: Sanjyal, Ankit
Published: (2025)
Multimodal Emotion Recognition via Causal-Diffusion Bridge (Affect-Diff)
by: Sanjyal, Ankit
Published: (2026)
by: Sanjyal, Ankit
Published: (2026)
In the Era of Prompt Learning with Vision-Language Models
by: Jha, Ankit
Published: (2024)
by: Jha, Ankit
Published: (2024)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
by: Jun-Seong, Kim, et al.
Published: (2026)
by: Jun-Seong, Kim, et al.
Published: (2026)
Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)
by: Lee, Jin-Seop, et al.
Published: (2025)
HDR-NSFF: High Dynamic Range Neural Scene Flow Fields
by: Dong-Yeon, Shin, et al.
Published: (2026)
by: Dong-Yeon, Shin, et al.
Published: (2026)
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
by: Shin, Jongmin, et al.
Published: (2026)
by: Shin, Jongmin, et al.
Published: (2026)
Putting the Object Back into Video Object Segmentation
by: Cheng, Ho Kei, et al.
Published: (2023)
by: Cheng, Ho Kei, et al.
Published: (2023)
RectifiedHR: High-Resolution Diffusion via Energy Profiling and Adaptive Guidance Scheduling
by: Sanjyal, Ankit
Published: (2025)
by: Sanjyal, Ankit
Published: (2025)
Continual-MEGA: A Large-scale Benchmark for Generalizable Continual Anomaly Detection
by: Lee, Geonu, et al.
Published: (2025)
by: Lee, Geonu, et al.
Published: (2025)
Factorized Multi-Resolution HashGrid for Efficient Neural Radiance Fields: Execution on Edge-Devices
by: Jun-Seong, Kim, et al.
Published: (2026)
by: Jun-Seong, Kim, et al.
Published: (2026)
Similar Items
-
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026) -
How can embedding models bind concepts?
by: Uselis, Arnas, et al.
Published: (2026) -
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026) -
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
by: Koishigarina, Darina, et al.
Published: (2025) -
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026)