Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Ontalvilla, Paula, Ormazabal, Aitor, Azkune, Gorka |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving the Efficiency of Visually Augmented Language Models
by: Ontalvilla, Paula, et al.
Published: (2024)
by: Ontalvilla, Paula, et al.
Published: (2024)
BertaQA: How Much Do Language Models Know About Local Culture?
by: Etxaniz, Julen, et al.
Published: (2024)
by: Etxaniz, Julen, et al.
Published: (2024)
When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively
by: Labruna, Tiziano, et al.
Published: (2024)
by: Labruna, Tiziano, et al.
Published: (2024)
Grounding Spatial Relations in Text-Only Language Models
by: Azkune, Gorka, et al.
Published: (2024)
by: Azkune, Gorka, et al.
Published: (2024)
Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
by: Arana, Lukas, et al.
Published: (2025)
by: Arana, Lukas, et al.
Published: (2025)
Vision-Language Models Struggle to Align Entities across Modalities
by: Alonso, Iñigo, et al.
Published: (2025)
by: Alonso, Iñigo, et al.
Published: (2025)
Adding simple structure at inference improves Vision-Language Compositionality
by: Miranda, Imanol, et al.
Published: (2025)
by: Miranda, Imanol, et al.
Published: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
by: Miranda, Imanol, et al.
Published: (2024)
by: Miranda, Imanol, et al.
Published: (2024)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
by: Miranda, Imanol, et al.
Published: (2026)
by: Miranda, Imanol, et al.
Published: (2026)
Conditioning LLMs to Generate Code-Switched Text
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
by: Attimonelli, Matteo, et al.
Published: (2026)
by: Attimonelli, Matteo, et al.
Published: (2026)
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
by: Wu, Qiyu, et al.
Published: (2025)
by: Wu, Qiyu, et al.
Published: (2025)
Latxa: An Open Language Model and Evaluation Suite for Basque
by: Etxaniz, Julen, et al.
Published: (2024)
by: Etxaniz, Julen, et al.
Published: (2024)
Can Language Models Compose Skills In-Context?
by: Liu, Zidong, et al.
Published: (2025)
by: Liu, Zidong, et al.
Published: (2025)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
How Do Language Models Compose Functions?
by: Khandelwal, Apoorv, et al.
Published: (2025)
by: Khandelwal, Apoorv, et al.
Published: (2025)
Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models
by: Van Doren, Madison, et al.
Published: (2025)
by: Van Doren, Madison, et al.
Published: (2025)
Do AI Models Perform Human-like Abstract Reasoning Across Modalities?
by: Beger, Claas, et al.
Published: (2025)
by: Beger, Claas, et al.
Published: (2025)
"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer?
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
by: Yuan, Lifan, et al.
Published: (2025)
by: Yuan, Lifan, et al.
Published: (2025)
Do LLMs and VLMs Share Neurons for Inference? Evidence and Mechanisms of Cross-Modal Transfer
by: Cui, Chenhang, et al.
Published: (2026)
by: Cui, Chenhang, et al.
Published: (2026)
Improving Explicit Spatial Relationships in Text-to-Image Generation through an Automatically Derived Dataset
by: Salaberria, Ander, et al.
Published: (2024)
by: Salaberria, Ander, et al.
Published: (2024)
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
by: Wang, Yucheng, et al.
Published: (2025)
by: Wang, Yucheng, et al.
Published: (2025)
Composing or Not Composing? Towards Distributional Construction Grammars
by: Blache, Philippe, et al.
Published: (2024)
by: Blache, Philippe, et al.
Published: (2024)
Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviews
by: Zou, Chenye, et al.
Published: (2025)
by: Zou, Chenye, et al.
Published: (2025)
Prompt Decorators: A Declarative and Composable Syntax for Reasoning, Formatting, and Control in LLMs
by: Heris, Mostapha Kalami
Published: (2025)
by: Heris, Mostapha Kalami
Published: (2025)
Skill-LLM: Repurposing General-Purpose LLMs for Skill Extraction
by: Herandi, Amirhossein, et al.
Published: (2024)
by: Herandi, Amirhossein, et al.
Published: (2024)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
by: Hua, Tianze, et al.
Published: (2025)
by: Hua, Tianze, et al.
Published: (2025)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents
by: Alvarez, Aitor Arronte, et al.
Published: (2026)
by: Alvarez, Aitor Arronte, et al.
Published: (2026)
Leveraging LLMs For Turkish Skill Extraction
by: İltüzer, Ezgi Arslan, et al.
Published: (2026)
by: İltüzer, Ezgi Arslan, et al.
Published: (2026)
Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality
by: Bu, Mengyu, et al.
Published: (2026)
by: Bu, Mengyu, et al.
Published: (2026)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
by: Maimon, Aviya, et al.
Published: (2025)
by: Maimon, Aviya, et al.
Published: (2025)
How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions
by: Kharchenko, Julia, et al.
Published: (2024)
by: Kharchenko, Julia, et al.
Published: (2024)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
by: Broomfield, Julius, et al.
Published: (2025)
by: Broomfield, Julius, et al.
Published: (2025)
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
by: Liu, Yujian, et al.
Published: (2026)
by: Liu, Yujian, et al.
Published: (2026)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
by: Sun, Kaiser, et al.
Published: (2026)
by: Sun, Kaiser, et al.
Published: (2026)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Similar Items
-
Improving the Efficiency of Visually Augmented Language Models
by: Ontalvilla, Paula, et al.
Published: (2024) -
BertaQA: How Much Do Language Models Know About Local Culture?
by: Etxaniz, Julen, et al.
Published: (2024) -
When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively
by: Labruna, Tiziano, et al.
Published: (2024) -
Grounding Spatial Relations in Text-Only Language Models
by: Azkune, Gorka, et al.
Published: (2024) -
Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
by: Arana, Lukas, et al.
Published: (2025)