Do Vision and Language Encoders Represent the World Similarly?
Fuente:
arXiv
Saved in:
| Main Authors: | Maniparambil, Mayug, Akshulakov, Raiymbek, Djilali, Yasser Abdelaziz Dahou, Narayan, Sanath, Seddik, Mohamed El Amine, Mangalam, Karttikeya, O'Connor, Noel E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024)
by: Narayan, Sanath, et al.
Published: (2024)
Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters
by: Vorster, Chris, et al.
Published: (2026)
by: Vorster, Chris, et al.
Published: (2026)
Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe
by: Vorster, Chris, et al.
Published: (2026)
by: Vorster, Chris, et al.
Published: (2026)
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
by: Hoehing, Nils, et al.
Published: (2025)
by: Hoehing, Nils, et al.
Published: (2025)
Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?
by: Hashmi, Anam, et al.
Published: (2026)
by: Hashmi, Anam, et al.
Published: (2026)
Ensemble Learning with Sparse Hypercolumns
by: Dietlmeier, Julia, et al.
Published: (2026)
by: Dietlmeier, Julia, et al.
Published: (2026)
Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
by: Sirotkin, Kirill, et al.
Published: (2024)
by: Sirotkin, Kirill, et al.
Published: (2024)
Falcon2-11B Technical Report
by: Malartic, Quentin, et al.
Published: (2024)
by: Malartic, Quentin, et al.
Published: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Adaptive Human Trajectory Prediction via Latent Corridors
by: Thakkar, Neerja, et al.
Published: (2023)
by: Thakkar, Neerja, et al.
Published: (2023)
Test-Time Adaptation with SaLIP: A Cascade of SAM and CLIP for Zero shot Medical Image Segmentation
by: Aleem, Sidra, et al.
Published: (2024)
by: Aleem, Sidra, et al.
Published: (2024)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
FASSILA: A Corpus for Algerian Dialect Fake News Detection and Sentiment Analysis
by: Abdedaiem, Amin, et al.
Published: (2024)
by: Abdedaiem, Amin, et al.
Published: (2024)
xT: Nested Tokenization for Larger Context in Large Images
by: Gupta, Ritwik, et al.
Published: (2024)
by: Gupta, Ritwik, et al.
Published: (2024)
Falcon Perception
by: Bevli, Aviraj, et al.
Published: (2026)
by: Bevli, Aviraj, et al.
Published: (2026)
How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models
by: Seddik, Mohamed El Amine
Published: (2026)
by: Seddik, Mohamed El Amine
Published: (2026)
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
by: Zhu, Zhanda, et al.
Published: (2025)
by: Zhu, Zhanda, et al.
Published: (2025)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
by: Younsi, Adam, et al.
Published: (2025)
by: Younsi, Adam, et al.
Published: (2025)
Dr$^2$Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
by: Zhao, Chen, et al.
Published: (2024)
by: Zhao, Chen, et al.
Published: (2024)
Where on Earth Do Users Say They Are?: Geo-Entity Linking for Noisy Multilingual User Input
by: Masis, Tessa, et al.
Published: (2024)
by: Masis, Tessa, et al.
Published: (2024)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
by: Lee, Nicholas, et al.
Published: (2024)
by: Lee, Nicholas, et al.
Published: (2024)
Synthetic Time Series for Anomaly Detection in Cloud Microservices
by: Allam, Mohamed, et al.
Published: (2024)
by: Allam, Mohamed, et al.
Published: (2024)
High-dimensional Learning with Noisy Labels
by: Firdoussi, Aymane El, et al.
Published: (2024)
by: Firdoussi, Aymane El, et al.
Published: (2024)
ESCape the ClassRoom
by: O'Connor, John
Published: (2024)
by: O'Connor, John
Published: (2024)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
by: Li, Boyi, et al.
Published: (2022)
by: Li, Boyi, et al.
Published: (2022)
A Monte Carlo Language Model Pipeline for Zero-Shot Sociopolitical Event Extraction
by: Cai, Erica, et al.
Published: (2023)
by: Cai, Erica, et al.
Published: (2023)
BERnaT: Basque Encoders for Representing Natural Textual Diversity
by: Azurmendi, Ekhi, et al.
Published: (2025)
by: Azurmendi, Ekhi, et al.
Published: (2025)
Beyond Traditional Single Object Tracking: A Survey
by: Abdelaziz, Omar, et al.
Published: (2024)
by: Abdelaziz, Omar, et al.
Published: (2024)
Do We Need Reformer for Vision? An Experimental Comparison with Vision Transformers
by: Bellaj, Ali El, et al.
Published: (2025)
by: Bellaj, Ali El, et al.
Published: (2025)
FSLI: An Interpretable Formal Semantic System for One-Dimensional Ordering Inference
by: Alkhairy, Maha, et al.
Published: (2025)
by: Alkhairy, Maha, et al.
Published: (2025)
Coordinates from Context: Using LLMs to Ground Complex Location References
by: Masis, Tessa, et al.
Published: (2025)
by: Masis, Tessa, et al.
Published: (2025)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
by: Yasser, Alaa, et al.
Published: (2026)
by: Yasser, Alaa, et al.
Published: (2026)
Understanding the Effect of Knowledge Graph Extraction Error on Downstream Graph Analyses: A Case Study on Affiliation Graphs
by: Cai, Erica, et al.
Published: (2025)
by: Cai, Erica, et al.
Published: (2025)
Formal Verification of the Safegcd Implementation
by: O'Connor, Russell, et al.
Published: (2025)
by: O'Connor, Russell, et al.
Published: (2025)
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction
by: Russell, Sam O'Connor, et al.
Published: (2025)
by: Russell, Sam O'Connor, et al.
Published: (2025)
Similar Items
-
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024) -
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024) -
Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters
by: Vorster, Chris, et al.
Published: (2026) -
Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe
by: Vorster, Chris, et al.
Published: (2026) -
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
by: Hoehing, Nils, et al.
Published: (2025)