The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Weichen, Peng, Ruiying, Zeng, Xin, Fang, Jianjie, Wang, Ziyou, Li, Kaiyuan, Dong, Heng, Li, Wei, Gao, Chen, Wang, Xin, Chen, Xinlei, Li, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding and Evaluating Hallucinations in 3D Visual Language Models
by: Peng, Ruiying, et al.
Published: (2025)
by: Peng, Ruiying, et al.
Published: (2025)
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
by: Fang, Jianjie, et al.
Published: (2026)
by: Fang, Jianjie, et al.
Published: (2026)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
by: Li, Jiaao, et al.
Published: (2025)
by: Li, Jiaao, et al.
Published: (2025)
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
by: Li, Kaiyuan, et al.
Published: (2025)
by: Li, Kaiyuan, et al.
Published: (2025)
Future Does Matter: Boosting 3D Object Detection with Temporal Motion Estimation in Point Cloud Sequences
by: Yu, Rui, et al.
Published: (2024)
by: Yu, Rui, et al.
Published: (2024)
Spatial Visibility and Temporal Dynamics: Revolutionizing Field of View Prediction in Adaptive Point Cloud Video Streaming
by: Li, Chen, et al.
Published: (2024)
by: Li, Chen, et al.
Published: (2024)
PointRegGPT: Boosting 3D Point Cloud Registration using Generative Point-Cloud Pairs for Training
by: Chen, Suyi, et al.
Published: (2024)
by: Chen, Suyi, et al.
Published: (2024)
Point Tree Transformer for Point Cloud Registration
by: Wang, Meiling, et al.
Published: (2024)
by: Wang, Meiling, et al.
Published: (2024)
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
by: Meng, Yu, et al.
Published: (2025)
by: Meng, Yu, et al.
Published: (2025)
Context-Aware Sentiment Forecasting via LLM-based Multi-Perspective Role-Playing Agents
by: Man, Fanhang, et al.
Published: (2025)
by: Man, Fanhang, et al.
Published: (2025)
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
by: Xu, Haotian, et al.
Published: (2026)
by: Xu, Haotian, et al.
Published: (2026)
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
by: Peng, Ruiying, et al.
Published: (2026)
by: Peng, Ruiying, et al.
Published: (2026)
Point-In-Context: Understanding Point Cloud via In-Context Learning
by: Liu, Mengyuan, et al.
Published: (2024)
by: Liu, Mengyuan, et al.
Published: (2024)
Adaptive and Iterative Point Cloud Denoising with Score‐Based Diffusion Model
by: Zhaonan Wang, et al.
Published: (2025)
by: Zhaonan Wang, et al.
Published: (2025)
Exploiting Topological Priors for Boosting Point Cloud Generation
by: Chen, Baiyuan
Published: (2024)
by: Chen, Baiyuan
Published: (2024)
AirScape: An Aerial Generative World Model with Motion Controllability
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
PointSmile: Point Self-supervised Learning via Curriculum Mutual Information
by: Li, Xin, et al.
Published: (2023)
by: Li, Xin, et al.
Published: (2023)
Adaptive and Iterative Point Cloud Denoising with Score-Based Diffusion Model
by: Wang, Zhaonan, et al.
Published: (2025)
by: Wang, Zhaonan, et al.
Published: (2025)
CK-MPM: A Compact-Kernel Material Point Method
by: Liu, Michael, et al.
Published: (2024)
by: Liu, Michael, et al.
Published: (2024)
Winding Clearness for Differentiable Point Cloud Optimization
by: Xiao, Dong, et al.
Published: (2024)
by: Xiao, Dong, et al.
Published: (2024)
Twin Deformable Point Convolutions for Point Cloud Semantic Segmentation in Remote Sensing Scenes
by: Mao, Yong-Qiang, et al.
Published: (2024)
by: Mao, Yong-Qiang, et al.
Published: (2024)
Mitigating Prior Shape Bias in Point Clouds via Differentiable Center Learning
by: Li, Zhe, et al.
Published: (2024)
by: Li, Zhe, et al.
Published: (2024)
PointSLAM++: Robust Dense Neural Gaussian Point Cloud-based SLAM
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Spiking Point Transformer for Point Cloud Classification
by: Wu, Peixi, et al.
Published: (2025)
by: Wu, Peixi, et al.
Published: (2025)
Progressive Frame Patching for FoV-based Point Cloud Video Streaming
by: Zong, Tongyu, et al.
Published: (2023)
by: Zong, Tongyu, et al.
Published: (2023)
CLIP-based Point Cloud Classification via Point Cloud to Image Translation
by: Ghose, Shuvozit, et al.
Published: (2024)
by: Ghose, Shuvozit, et al.
Published: (2024)
PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Hierarchical Attention Networks for Lossless Point Cloud Attribute Compression
by: Chen, Yueru, et al.
Published: (2025)
by: Chen, Yueru, et al.
Published: (2025)
BADet: Boundary-Aware 3D Object Detection from Point Clouds
by: Qian, Rui, et al.
Published: (2021)
by: Qian, Rui, et al.
Published: (2021)
PointSCNet: Point Cloud Structure and Correlation Learning Based on Space Filling Curve-Guided Sampling
by: Chen, Xingye, et al.
Published: (2022)
by: Chen, Xingye, et al.
Published: (2022)
Masked Generative Extractor for Synergistic Representation and 3D Generation of Point Clouds
by: Zeng, Hongliang, et al.
Published: (2024)
by: Zeng, Hongliang, et al.
Published: (2024)
Joint Point Cloud Upsampling and Cleaning with Octree-based CNNs
by: Li, Jihe, et al.
Published: (2024)
by: Li, Jihe, et al.
Published: (2024)
Efficient Spiking Point Mamba for Point Cloud Analysis
by: Wu, Peixi, et al.
Published: (2025)
by: Wu, Peixi, et al.
Published: (2025)
PCSTracker: Long-Term Scene Flow Estimation for Point Cloud Sequences
by: Lin, Min, et al.
Published: (2026)
by: Lin, Min, et al.
Published: (2026)
Similar Items
-
Understanding and Evaluating Hallucinations in 3D Visual Language Models
by: Peng, Ruiying, et al.
Published: (2025) -
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
by: Zhao, Baining, et al.
Published: (2025) -
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025) -
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
by: Zhao, Baining, et al.
Published: (2025) -
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
by: Fang, Jianjie, et al.
Published: (2026)