Multimodal Language Models Cannot Spot Spatial Inconsistencies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khangaonkar, Om, Rad, Hadi J., Pirsiavash, Hamed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
gen2seg: Generative Models Enable Generalizable Instance Segmentation
von: Khangaonkar, Om, et al.
Veröffentlicht: (2025)
von: Khangaonkar, Om, et al.
Veröffentlicht: (2025)
One Category One Prompt: Dataset Distillation using Diffusion Models
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
von: Kumar, Divake, et al.
Veröffentlicht: (2026)
von: Kumar, Divake, et al.
Veröffentlicht: (2026)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
MultiDelete for Multimodal Machine Unlearning
von: Cheng, Jiali, et al.
Veröffentlicht: (2023)
von: Cheng, Jiali, et al.
Veröffentlicht: (2023)
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
von: Hamed, Omar, et al.
Veröffentlicht: (2024)
von: Hamed, Omar, et al.
Veröffentlicht: (2024)
LaVy: Vietnamese Multimodal Large Language Model
von: Tran, Chi, et al.
Veröffentlicht: (2024)
von: Tran, Chi, et al.
Veröffentlicht: (2024)
Multimodal Latent Language Modeling with Next-Token Diffusion
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
Improving Multimodal Large Language Models Using Continual Learning
von: Srivastava, Shikhar, et al.
Veröffentlicht: (2024)
von: Srivastava, Shikhar, et al.
Veröffentlicht: (2024)
On Domain-Adaptive Post-Training for Multimodal Large Language Models
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
von: Ganesan, Mugilan, et al.
Veröffentlicht: (2025)
von: Ganesan, Mugilan, et al.
Veröffentlicht: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
Unified Control for Inference-Time Guidance of Denoising Diffusion Models
von: Goyal, Maurya, et al.
Veröffentlicht: (2025)
von: Goyal, Maurya, et al.
Veröffentlicht: (2025)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
NOLA: Compressing LoRA using Linear Combination of Random Basis
von: Koohpayegani, Soroush Abbasi, et al.
Veröffentlicht: (2023)
von: Koohpayegani, Soroush Abbasi, et al.
Veröffentlicht: (2023)
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
GeNIe: Generative Hard Negative Images Through Diffusion
von: Koohpayegani, Soroush Abbasi, et al.
Veröffentlicht: (2023)
von: Koohpayegani, Soroush Abbasi, et al.
Veröffentlicht: (2023)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
BECLR: Batch Enhanced Contrastive Few-Shot Learning
von: Poulakakis-Daktylidis, Stylianos, et al.
Veröffentlicht: (2024)
von: Poulakakis-Daktylidis, Stylianos, et al.
Veröffentlicht: (2024)
Distributed LLMs and Multimodal Large Language Models: A Survey on Advances, Challenges, and Future Directions
von: Amini, Hadi, et al.
Veröffentlicht: (2025)
von: Amini, Hadi, et al.
Veröffentlicht: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
How to Merge Your Multimodal Models Over Time?
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2024)
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2024)
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)
von: Bai, Lei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
gen2seg: Generative Models Enable Generalizable Instance Segmentation
von: Khangaonkar, Om, et al.
Veröffentlicht: (2025) -
One Category One Prompt: Dataset Distillation using Diffusion Models
von: Abbasi, Ali, et al.
Veröffentlicht: (2024) -
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
von: Kumar, Divake, et al.
Veröffentlicht: (2026) -
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023) -
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)