Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
Fuente:
arXiv
Saved in:
| Main Authors: | Alam, Nahid, Murali, Leema Krishna, Bharadwaj, Siddhant, Liu, Patrick, Chung, Timothy, Sharma, Drishti, A., Akshata, Kiran, Kranthi, Tam, Wesley, Vegesna, Bala Krishna S |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Spatial Blindspot of Vision-Language Models
by: Alam, Nahid, et al.
Published: (2026)
by: Alam, Nahid, et al.
Published: (2026)
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
by: Kanjula, Karthik Reddy, et al.
Published: (2025)
by: Kanjula, Karthik Reddy, et al.
Published: (2025)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
by: Zeer, Ahmed, et al.
Published: (2024)
by: Zeer, Ahmed, et al.
Published: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
by: Yan, Dawei, et al.
Published: (2024)
by: Yan, Dawei, et al.
Published: (2024)
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
by: Zafar, Anas, et al.
Published: (2026)
by: Zafar, Anas, et al.
Published: (2026)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
by: Seyfioglu, Mehmet Saygin, et al.
Published: (2023)
by: Seyfioglu, Mehmet Saygin, et al.
Published: (2023)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
by: Lou, Haoran, et al.
Published: (2025)
by: Lou, Haoran, et al.
Published: (2025)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
by: Sun, Xuefei, et al.
Published: (2025)
by: Sun, Xuefei, et al.
Published: (2025)
LLaVA-Critic: Learning to Evaluate Multimodal Models
by: Xiong, Tianyi, et al.
Published: (2024)
by: Xiong, Tianyi, et al.
Published: (2024)
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
by: Chay-intr, T., et al.
Published: (2025)
by: Chay-intr, T., et al.
Published: (2025)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
by: Andersland, Michael
Published: (2024)
by: Andersland, Michael
Published: (2024)
Unconventional optical response in monolayer graphene upon dominant intraband scattering
by: Saha, Palash, et al.
Published: (2023)
by: Saha, Palash, et al.
Published: (2023)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
by: Shen, Leqi, et al.
Published: (2025)
by: Shen, Leqi, et al.
Published: (2025)
Can Sound Replace Vision in LLaVA With Token Substitution?
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
by: Ding, Zhicheng, et al.
Published: (2024)
by: Ding, Zhicheng, et al.
Published: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
by: Li, Bo, et al.
Published: (2024)
by: Li, Bo, et al.
Published: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
by: Sun, Guohao, et al.
Published: (2024)
by: Sun, Guohao, et al.
Published: (2024)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
by: Caffagni, Davide, et al.
Published: (2024)
by: Caffagni, Davide, et al.
Published: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
by: Liang, Han, et al.
Published: (2024)
by: Liang, Han, et al.
Published: (2024)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
by: Guo, Xuechen, et al.
Published: (2024)
by: Guo, Xuechen, et al.
Published: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
by: An, Ruichuan, et al.
Published: (2024)
by: An, Ruichuan, et al.
Published: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
by: An, Ruichuan, et al.
Published: (2025)
by: An, Ruichuan, et al.
Published: (2025)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
by: Inal, Gokce, et al.
Published: (2026)
by: Inal, Gokce, et al.
Published: (2026)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
by: Shin, Dongjae, et al.
Published: (2024)
by: Shin, Dongjae, et al.
Published: (2024)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
by: Sun, Shenghuan, et al.
Published: (2024)
by: Sun, Shenghuan, et al.
Published: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
by: Yoon, Kyung-Yoon, et al.
Published: (2025)
by: Yoon, Kyung-Yoon, et al.
Published: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
by: Zhu, Yichen, et al.
Published: (2024)
by: Zhu, Yichen, et al.
Published: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
by: Shi, Wenhao, et al.
Published: (2024)
by: Shi, Wenhao, et al.
Published: (2024)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
by: Lin, Bin, et al.
Published: (2023)
by: Lin, Bin, et al.
Published: (2023)
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
by: Foutter, Matthew, et al.
Published: (2024)
by: Foutter, Matthew, et al.
Published: (2024)
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
by: Zhang, Yi-Fan, et al.
Published: (2024)
by: Zhang, Yi-Fan, et al.
Published: (2024)
Similar Items
-
The Spatial Blindspot of Vision-Language Models
by: Alam, Nahid, et al.
Published: (2026) -
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
by: Kanjula, Karthik Reddy, et al.
Published: (2025) -
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
by: Zeer, Ahmed, et al.
Published: (2024) -
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
by: Yan, Dawei, et al.
Published: (2024) -
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
by: Zafar, Anas, et al.
Published: (2026)