Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Xuefei, Albin, Doncey, Mauceri, Cecilia, Woods, Dusty, Heckman, Christoffer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CU-Multi: A Dataset for Multi-Robot Data Association
by: Albin, Doncey, et al.
Published: (2025)
by: Albin, Doncey, et al.
Published: (2025)
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
by: Sun, Xuefei, et al.
Published: (2026)
by: Sun, Xuefei, et al.
Published: (2026)
CogExplore: Contextual Exploration with Language-Encoded Environment Representations
by: Biggie, Harel, et al.
Published: (2024)
by: Biggie, Harel, et al.
Published: (2024)
SceneSense: Diffusion Models for 3D Occupancy Synthesis from Partial Observation
by: Reed, Alec, et al.
Published: (2024)
by: Reed, Alec, et al.
Published: (2024)
Cost-Effective Radar Sensors for Field-Based Water Level Monitoring with Sub-Centimeter Accuracy
by: Zavei-Boroda, Anna, et al.
Published: (2026)
by: Zavei-Boroda, Anna, et al.
Published: (2026)
CU-Multi: A Dataset for Multi-Robot Collaborative Perception
by: Albin, Doncey, et al.
Published: (2025)
by: Albin, Doncey, et al.
Published: (2025)
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
by: Foutter, Matthew, et al.
Published: (2024)
by: Foutter, Matthew, et al.
Published: (2024)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
by: Lou, Haoran, et al.
Published: (2025)
by: Lou, Haoran, et al.
Published: (2025)
Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces
by: Payandeh, Amirreza, et al.
Published: (2024)
by: Payandeh, Amirreza, et al.
Published: (2024)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Radar-Based Localization For Autonomous Ground Vehicles In Suburban Neighborhoods
by: Kramer, Andrew J., et al.
Published: (2024)
by: Kramer, Andrew J., et al.
Published: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
by: Zeer, Ahmed, et al.
Published: (2024)
by: Zeer, Ahmed, et al.
Published: (2024)
Weather-Robust Scene Semantics with Vision-Aligned 4D Radar
by: Hamilton, Kali, et al.
Published: (2026)
by: Hamilton, Kali, et al.
Published: (2026)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
by: Wang, Jingyi, et al.
Published: (2024)
by: Wang, Jingyi, et al.
Published: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
Foundation Models for Rapid Autonomy Validation
by: Farid, Alec, et al.
Published: (2024)
by: Farid, Alec, et al.
Published: (2024)
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
by: Chaubey, Ashutosh, et al.
Published: (2025)
by: Chaubey, Ashutosh, et al.
Published: (2025)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
RF-Modulated Adaptive Communication Improves Multi-Agent Robotic Exploration
by: Achey, Lorin, et al.
Published: (2026)
by: Achey, Lorin, et al.
Published: (2026)
Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
by: Alam, Nahid, et al.
Published: (2026)
by: Alam, Nahid, et al.
Published: (2026)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
by: Yan, Dawei, et al.
Published: (2024)
by: Yan, Dawei, et al.
Published: (2024)
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations
by: Xu, Mingjie, et al.
Published: (2024)
by: Xu, Mingjie, et al.
Published: (2024)
Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding
by: Patratskiy, Maxim A., et al.
Published: (2025)
by: Patratskiy, Maxim A., et al.
Published: (2025)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
by: Sun, Guohao, et al.
Published: (2024)
by: Sun, Guohao, et al.
Published: (2024)
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
by: Ma, Xianzheng, et al.
Published: (2026)
by: Ma, Xianzheng, et al.
Published: (2026)
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
by: Jin, Juseong, et al.
Published: (2024)
by: Jin, Juseong, et al.
Published: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
by: Liang, Han, et al.
Published: (2024)
by: Liang, Han, et al.
Published: (2024)
Belief Consistency Between Foundation-Model Evidence and Geometric Perception in Persistent Robotic Maps
by: Heckman, Christoffer, et al.
Published: (2026)
by: Heckman, Christoffer, et al.
Published: (2026)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
by: Lan, Zhibin, et al.
Published: (2024)
by: Lan, Zhibin, et al.
Published: (2024)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
by: Sun, Shenghuan, et al.
Published: (2024)
by: Sun, Shenghuan, et al.
Published: (2024)
Online Diffusion-Based 3D Occupancy Prediction at the Frontier with Probabilistic Map Reconciliation
by: Reed, Alec, et al.
Published: (2024)
by: Reed, Alec, et al.
Published: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
by: Shi, Wenhao, et al.
Published: (2024)
by: Shi, Wenhao, et al.
Published: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
by: Ye, Xubing, et al.
Published: (2024)
by: Ye, Xubing, et al.
Published: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
PA-LLaVA: A Large Language-Vision Assistant for Human Pathology Image Understanding
by: Dai, Dawei, et al.
Published: (2024)
by: Dai, Dawei, et al.
Published: (2024)
Similar Items
-
CU-Multi: A Dataset for Multi-Robot Data Association
by: Albin, Doncey, et al.
Published: (2025) -
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
by: Sun, Xuefei, et al.
Published: (2026) -
CogExplore: Contextual Exploration with Language-Encoded Environment Representations
by: Biggie, Harel, et al.
Published: (2024) -
SceneSense: Diffusion Models for 3D Occupancy Synthesis from Partial Observation
by: Reed, Alec, et al.
Published: (2024) -
Cost-Effective Radar Sensors for Field-Based Water Level Monitoring with Sub-Centimeter Accuracy
by: Zavei-Boroda, Anna, et al.
Published: (2026)