Vision-Language Models Struggle to Align Entities across Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Alonso, Iñigo, Azkune, Gorka, Salaberria, Ander, Barnes, Jeremy, de Lacalle, Oier Lopez |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounding Spatial Relations in Text-Only Language Models
by: Azkune, Gorka, et al.
Published: (2024)
by: Azkune, Gorka, et al.
Published: (2024)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
by: Miranda, Imanol, et al.
Published: (2026)
by: Miranda, Imanol, et al.
Published: (2026)
Adding simple structure at inference improves Vision-Language Compositionality
by: Miranda, Imanol, et al.
Published: (2025)
by: Miranda, Imanol, et al.
Published: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
by: Miranda, Imanol, et al.
Published: (2024)
by: Miranda, Imanol, et al.
Published: (2024)
Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
by: Arana, Lukas, et al.
Published: (2025)
by: Arana, Lukas, et al.
Published: (2025)
Improving Explicit Spatial Relationships in Text-to-Image Generation through an Automatically Derived Dataset
by: Salaberria, Ander, et al.
Published: (2024)
by: Salaberria, Ander, et al.
Published: (2024)
BertaQA: How Much Do Language Models Know About Local Culture?
by: Etxaniz, Julen, et al.
Published: (2024)
by: Etxaniz, Julen, et al.
Published: (2024)
Reasoning over Object Descriptions Improves Coreference Resolution in Task-Based Dialogue Systems
by: Ijurco, Oier, et al.
Published: (2026)
by: Ijurco, Oier, et al.
Published: (2026)
When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively
by: Labruna, Tiziano, et al.
Published: (2024)
by: Labruna, Tiziano, et al.
Published: (2024)
Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
by: Ontalvilla, Paula, et al.
Published: (2025)
by: Ontalvilla, Paula, et al.
Published: (2025)
Improving the Efficiency of Visually Augmented Language Models
by: Ontalvilla, Paula, et al.
Published: (2024)
by: Ontalvilla, Paula, et al.
Published: (2024)
Euskarazko lehen C1 ebaluatzaile automatikoa
by: Azurmendi, Ekhi, et al.
Published: (2025)
by: Azurmendi, Ekhi, et al.
Published: (2025)
Automatic Essay Scoring and Feedback Generation in Basque Language Learning
by: Azurmendi, Ekhi, et al.
Published: (2025)
by: Azurmendi, Ekhi, et al.
Published: (2025)
Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction
by: Zubillaga, Mikel, et al.
Published: (2026)
by: Zubillaga, Mikel, et al.
Published: (2026)
Event Extraction in Basque: Typologically motivated Cross-Lingual Transfer-Learning Analysis
by: Zubillaga, Mikel, et al.
Published: (2024)
by: Zubillaga, Mikel, et al.
Published: (2024)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
by: Cohen, Ido, et al.
Published: (2024)
by: Cohen, Ido, et al.
Published: (2024)
GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction
by: Sainz, Oscar, et al.
Published: (2023)
by: Sainz, Oscar, et al.
Published: (2023)
Conditioning LLMs to Generate Code-Switched Text
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
Deriving MOND from f(T) teleparallel gravity: a selected interpolation function and acceleration scale
by: Ander, Azkune
Published: (2026)
by: Ander, Azkune
Published: (2026)
Large Language Models Struggle in Token-Level Clinical Named Entity Recognition
by: Lu, Qiuhao, et al.
Published: (2024)
by: Lu, Qiuhao, et al.
Published: (2024)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
by: Zhou, Yiyang, et al.
Published: (2024)
by: Zhou, Yiyang, et al.
Published: (2024)
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
by: Alonso, Iñigo, et al.
Published: (2024)
by: Alonso, Iñigo, et al.
Published: (2024)
Signs of Struggle: Spotting Cognitive Distortions across Language and Register
by: Kuber, Abhishek, et al.
Published: (2025)
by: Kuber, Abhishek, et al.
Published: (2025)
Neural Correlates of Language Models Are Specific to Human Language
by: Parra, Iñigo
Published: (2025)
by: Parra, Iñigo
Published: (2025)
LLM-Align: Utilizing Large Language Models for Entity Alignment in Knowledge Graphs
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
Vision-Language Models Align with Human Neural Representations in Concept Processing
by: Bavaresco, Anna, et al.
Published: (2024)
by: Bavaresco, Anna, et al.
Published: (2024)
Progressively Modality Freezing for Multi-Modal Entity Alignment
by: Huang, Yani, et al.
Published: (2024)
by: Huang, Yani, et al.
Published: (2024)
Source-Modality Monitoring in Vision-Language Models
by: Hua, Etha Tianze, et al.
Published: (2026)
by: Hua, Etha Tianze, et al.
Published: (2026)
What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning
by: Weng, Zhaotian, et al.
Published: (2025)
by: Weng, Zhaotian, et al.
Published: (2025)
Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity Recognition
by: Yang, Yuming, et al.
Published: (2024)
by: Yang, Yuming, et al.
Published: (2024)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
by: Deng, Ken, et al.
Published: (2026)
by: Deng, Ken, et al.
Published: (2026)
Large Language Models Struggle with Unreasonability in Math Problems
by: Ma, Jingyuan, et al.
Published: (2024)
by: Ma, Jingyuan, et al.
Published: (2024)
Automatic Logical Forms improve fidelity in Table-to-Text generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
English Prompts are Better for NLI-based Zero-Shot Emotion Classification than Target-Language Prompts
by: Bareiß, Patrick, et al.
Published: (2024)
by: Bareiß, Patrick, et al.
Published: (2024)
On Entity Identification in Language Models
by: Sakata, Masaki, et al.
Published: (2025)
by: Sakata, Masaki, et al.
Published: (2025)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
by: Lucy, Li, et al.
Published: (2026)
by: Lucy, Li, et al.
Published: (2026)
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
by: Cohen, Vanya, et al.
Published: (2025)
by: Cohen, Vanya, et al.
Published: (2025)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Similar Items
-
Grounding Spatial Relations in Text-Only Language Models
by: Azkune, Gorka, et al.
Published: (2024) -
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
by: Miranda, Imanol, et al.
Published: (2026) -
Adding simple structure at inference improves Vision-Language Compositionality
by: Miranda, Imanol, et al.
Published: (2025) -
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
by: Miranda, Imanol, et al.
Published: (2024) -
Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
by: Arana, Lukas, et al.
Published: (2025)