LLMs Can Compensate for Deficiencies in Visual Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Takishita, Sho, Gala, Jay, Mohamed, Abdelrahman, Inui, Kentaro, Kementchedjhieva, Yova |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025)
by: Mohamed, Abdelrahman, et al.
Published: (2025)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
by: Li, Jiaang, et al.
Published: (2023)
by: Li, Jiaang, et al.
Published: (2023)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
by: Shoer, Belal, et al.
Published: (2025)
by: Shoer, Belal, et al.
Published: (2025)
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
by: Huzaifa, Muhammad, et al.
Published: (2024)
by: Huzaifa, Muhammad, et al.
Published: (2024)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
by: Salazar, Israfel, et al.
Published: (2025)
by: Salazar, Israfel, et al.
Published: (2025)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025)
by: Khayatan, Pegah, et al.
Published: (2025)
Can VLMs Recall Factual Associations From Visual References?
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Can Vision-Language Models Solve Visual Math Equations?
by: Choudhury, Monjoy Narayan, et al.
Published: (2025)
by: Choudhury, Monjoy Narayan, et al.
Published: (2025)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
by: Gado, Mohamed, et al.
Published: (2025)
by: Gado, Mohamed, et al.
Published: (2025)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
by: Vaishnav, Mohit, et al.
Published: (2026)
by: Vaishnav, Mohit, et al.
Published: (2026)
Give me a hint: Can LLMs take a hint to solve math problems?
by: Agrawal, Vansh, et al.
Published: (2024)
by: Agrawal, Vansh, et al.
Published: (2024)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
by: Bharti, Shubham, et al.
Published: (2024)
by: Bharti, Shubham, et al.
Published: (2024)
Can Rule-Based Insights Enhance LLMs for Radiology Report Classification? Introducing the RadPrompt Methodology
by: Fytas, Panagiotis, et al.
Published: (2024)
by: Fytas, Panagiotis, et al.
Published: (2024)
Multimodal LLMs Struggle with Basic Visual Network Analysis: a VNA Benchmark
by: Williams, Evan M., et al.
Published: (2024)
by: Williams, Evan M., et al.
Published: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
by: Jiang, Ruixiang, et al.
Published: (2025)
by: Jiang, Ruixiang, et al.
Published: (2025)
Compensating Distribution Drifts in Class-incremental Learning of Pre-trained Vision Transformers
by: Rao, Xuan, et al.
Published: (2025)
by: Rao, Xuan, et al.
Published: (2025)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
by: Chen, Danlu, et al.
Published: (2024)
by: Chen, Danlu, et al.
Published: (2024)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
by: Wang, Junling, et al.
Published: (2026)
by: Wang, Junling, et al.
Published: (2026)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
by: Zhou, Wenrui, et al.
Published: (2025)
by: Zhou, Wenrui, et al.
Published: (2025)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
by: Kawasaki, Haruka, et al.
Published: (2026)
by: Kawasaki, Haruka, et al.
Published: (2026)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
by: Yuan, Fan, et al.
Published: (2025)
by: Yuan, Fan, et al.
Published: (2025)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
by: Yu, Zhuoran, et al.
Published: (2025)
by: Yu, Zhuoran, et al.
Published: (2025)
Multimodal Large Language Models to Support Real-World Fact-Checking
by: Geng, Jiahui, et al.
Published: (2024)
by: Geng, Jiahui, et al.
Published: (2024)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
by: Lai, Zhengzhao, et al.
Published: (2025)
by: Lai, Zhengzhao, et al.
Published: (2025)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
by: Cocchi, Federico, et al.
Published: (2025)
by: Cocchi, Federico, et al.
Published: (2025)
Can Visual Encoder Learn to See Arrows?
by: Terashita, Naoyuki, et al.
Published: (2025)
by: Terashita, Naoyuki, et al.
Published: (2025)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
by: Cocchi, Federico, et al.
Published: (2024)
by: Cocchi, Federico, et al.
Published: (2024)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
by: Han, Jiaming, et al.
Published: (2025)
by: Han, Jiaming, et al.
Published: (2025)
Similar Items
-
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025) -
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
by: Li, Jiaang, et al.
Published: (2023) -
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025) -
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
by: Shoer, Belal, et al.
Published: (2025) -
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
by: Huzaifa, Muhammad, et al.
Published: (2024)