Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rajabi, Navid, Kosecka, Jana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
von: Rajabi, Navid, et al.
Veröffentlicht: (2025)
von: Rajabi, Navid, et al.
Veröffentlicht: (2025)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
von: Fayyazsanavi, Pooya, et al.
Veröffentlicht: (2024)
von: Fayyazsanavi, Pooya, et al.
Veröffentlicht: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
Multi-Modal Hallucination Control by Visual Information Grounding
von: Favero, Alessandro, et al.
Veröffentlicht: (2024)
von: Favero, Alessandro, et al.
Veröffentlicht: (2024)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
Composition-Grounded Data Synthesis for Visual Reasoning
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2026)
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2026)
Towards Visual Text Grounding of Multimodal Large Language Model
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
von: Ding, Yi, et al.
Veröffentlicht: (2025)
von: Ding, Yi, et al.
Veröffentlicht: (2025)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
Towards Statistical Factuality Guarantee for Large Vision-Language Models
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
von: Li, Sijie, et al.
Veröffentlicht: (2026)
von: Li, Sijie, et al.
Veröffentlicht: (2026)
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2026)
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2026)
Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
von: Lian, Shijie, et al.
Veröffentlicht: (2025)
von: Lian, Shijie, et al.
Veröffentlicht: (2025)
Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
von: Hong, Rui, et al.
Veröffentlicht: (2026)
von: Hong, Rui, et al.
Veröffentlicht: (2026)
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
von: Li, Xue, et al.
Veröffentlicht: (2025)
von: Li, Xue, et al.
Veröffentlicht: (2025)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
von: Pathak, Surendra, et al.
Veröffentlicht: (2026)
von: Pathak, Surendra, et al.
Veröffentlicht: (2026)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
von: Filvantorkaman, Melika, et al.
Veröffentlicht: (2026)
von: Filvantorkaman, Melika, et al.
Veröffentlicht: (2026)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
Why are Visually-Grounded Language Models Bad at Image Classification?
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
Structured Spatial Reasoning with Open Vocabulary Object Detectors
von: Nejatishahidin, Negar, et al.
Veröffentlicht: (2024)
von: Nejatishahidin, Negar, et al.
Veröffentlicht: (2024)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models
von: Zhang, Letian, et al.
Veröffentlicht: (2023)
von: Zhang, Letian, et al.
Veröffentlicht: (2023)
Learning from Synthetic Data for Visual Grounding
von: He, Ruozhen, et al.
Veröffentlicht: (2024)
von: He, Ruozhen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
von: Rajabi, Navid, et al.
Veröffentlicht: (2024) -
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
von: Rajabi, Navid, et al.
Veröffentlicht: (2024) -
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
von: Rajabi, Navid, et al.
Veröffentlicht: (2025) -
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
von: Fayyazsanavi, Pooya, et al.
Veröffentlicht: (2024) -
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)