VoxRep: Enhancing 3D Spatial Understanding in 2D Vision-Language Models via Voxel Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Dao, Alan, Buppodom, Norapat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
Lucy: edgerunning agentic web search on mobile with machine generated task vectors
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
VoxNeuS: Enhancing Voxel-Based Neural Surface Reconstruction via Gradient Interpolation
by: Liu, Sidun, et al.
Published: (2024)
by: Liu, Sidun, et al.
Published: (2024)
SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection
by: Son, Hyeongseok, et al.
Published: (2025)
by: Son, Hyeongseok, et al.
Published: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)
by: Zheng, Duo, et al.
Published: (2024)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
VoxNeRF: Bridging Voxel Representation and Neural Radiance Fields for Enhanced Indoor View Synthesis
by: Wang, Sen, et al.
Published: (2023)
by: Wang, Sen, et al.
Published: (2023)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
by: Sun, Haowen, et al.
Published: (2026)
by: Sun, Haowen, et al.
Published: (2026)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
by: Daxberger, Erik, et al.
Published: (2025)
by: Daxberger, Erik, et al.
Published: (2025)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026)
by: Liang, Huizhi, et al.
Published: (2026)
Vision-Language Models Do Not Understand Negation
by: Alhamoud, Kumail, et al.
Published: (2025)
by: Alhamoud, Kumail, et al.
Published: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
How to Utilize Complementary Vision-Text Information for 2D Structure Understanding
by: Dong, Jiancheng, et al.
Published: (2026)
by: Dong, Jiancheng, et al.
Published: (2026)
E3D-GPT: Enhanced 3D Visual Foundation for Medical Vision-Language Model
by: Lai, Haoran, et al.
Published: (2024)
by: Lai, Haoran, et al.
Published: (2024)
Investigating Spatial Attention Bias in Vision-Language Models
by: Chaudhary, Aryan, et al.
Published: (2025)
by: Chaudhary, Aryan, et al.
Published: (2025)
Do Vision-Language Models Really Understand Visual Language?
by: Hou, Yifan, et al.
Published: (2024)
by: Hou, Yifan, et al.
Published: (2024)
VoxelTrack: Exploring Voxel Representation for 3D Point Cloud Object Tracking
by: Lu, Yuxuan, et al.
Published: (2024)
by: Lu, Yuxuan, et al.
Published: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
by: Hua, Jiacheng, et al.
Published: (2026)
by: Hua, Jiacheng, et al.
Published: (2026)
Do Vision-Language Models Understand Visual Persuasiveness?
by: Park, Gyuwon
Published: (2025)
by: Park, Gyuwon
Published: (2025)
Do Vision-Language Models Understand Compound Nouns?
by: Kumar, Sonal, et al.
Published: (2024)
by: Kumar, Sonal, et al.
Published: (2024)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
by: Stogiannidis, Ilias, et al.
Published: (2025)
by: Stogiannidis, Ilias, et al.
Published: (2025)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
by: Zhu, Yingjie, et al.
Published: (2025)
by: Zhu, Yingjie, et al.
Published: (2025)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
by: Xu, Runsen, et al.
Published: (2025)
by: Xu, Runsen, et al.
Published: (2025)
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
by: Wei, Canshi
Published: (2024)
by: Wei, Canshi
Published: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
PUMGPT: A Large Vision-Language Model for Product Understanding
by: Xue, Wei, et al.
Published: (2023)
by: Xue, Wei, et al.
Published: (2023)
Toward Interactive Regional Understanding in Vision-Large Language Models
by: Lee, Jungbeom, et al.
Published: (2024)
by: Lee, Jungbeom, et al.
Published: (2024)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
VoxCor: Training-Free Volumetric Features for Multimodal Voxel Correspondence
by: Tombak, Guney, et al.
Published: (2026)
by: Tombak, Guney, et al.
Published: (2026)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
by: Du, Mengfei, et al.
Published: (2024)
by: Du, Mengfei, et al.
Published: (2024)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
by: Zhan, Shaoxiong, et al.
Published: (2026)
by: Zhan, Shaoxiong, et al.
Published: (2026)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
by: Ahmed, Mahmoud, et al.
Published: (2025)
by: Ahmed, Mahmoud, et al.
Published: (2025)
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
by: Huang, Sheng-Yu, et al.
Published: (2026)
by: Huang, Sheng-Yu, et al.
Published: (2026)
Granular Privacy Control for Geolocation with Vision Language Models
by: Mendes, Ethan, et al.
Published: (2024)
by: Mendes, Ethan, et al.
Published: (2024)
Transcrib3D: 3D Referring Expression Resolution through Large Language Models
by: Fang, Jiading, et al.
Published: (2024)
by: Fang, Jiading, et al.
Published: (2024)
Language-Image Models with 3D Understanding
by: Cho, Jang Hyun, et al.
Published: (2024)
by: Cho, Jang Hyun, et al.
Published: (2024)
Similar Items
-
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024) -
Lucy: edgerunning agentic web search on mobile with machine generated task vectors
by: Dao, Alan, et al.
Published: (2025) -
VoxNeuS: Enhancing Voxel-Based Neural Surface Reconstruction via Gradient Interpolation
by: Liu, Sidun, et al.
Published: (2024) -
SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection
by: Son, Hyeongseok, et al.
Published: (2025) -
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)