3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Dingning, Wang, Cheng, Gao, Peng, Zhang, Renrui, Ma, Xinzhu, Meng, Yuan, Wang, Zhihui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
by: Liu, Dingning, et al.
Published: (2024)
by: Liu, Dingning, et al.
Published: (2024)
Learning Geometry-Guided Depth via Projective Modeling for Monocular 3D Object Detection
by: Zhang, Yinmin, et al.
Published: (2021)
by: Zhang, Yinmin, et al.
Published: (2021)
Referring Video Object Segmentation with Cross-Modality Proxy Queries
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth Completion
by: Chen, Shenglun, et al.
Published: (2025)
by: Chen, Shenglun, et al.
Published: (2025)
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
by: Zhang, Xuying, et al.
Published: (2024)
by: Zhang, Xuying, et al.
Published: (2024)
3D Object Detection from Images for Autonomous Driving: A Survey
by: Ma, Xinzhu, et al.
Published: (2022)
by: Ma, Xinzhu, et al.
Published: (2022)
R2G: Reasoning to Ground in 3D Scenes
by: Li, Yixuan, et al.
Published: (2024)
by: Li, Yixuan, et al.
Published: (2024)
Reasoning Matters for 3D Visual Grounding
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Point2Primitive: CAD Reconstruction from Point Cloud by Direct Primitive Prediction
by: Ma, Xinzhu, et al.
Published: (2025)
by: Ma, Xinzhu, et al.
Published: (2025)
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
by: Chen, Wenyue, et al.
Published: (2026)
by: Chen, Wenyue, et al.
Published: (2026)
Towards Efficient and Intelligent Laser Weeding: Method and Dataset for Weed Stem Detection
by: Liu, Dingning, et al.
Published: (2025)
by: Liu, Dingning, et al.
Published: (2025)
Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts
by: Cheng, Xinhua, et al.
Published: (2023)
by: Cheng, Xinhua, et al.
Published: (2023)
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding
by: Zheng, Henry, et al.
Published: (2026)
by: Zheng, Henry, et al.
Published: (2026)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
by: Yang, Jiahui, et al.
Published: (2024)
by: Yang, Jiahui, et al.
Published: (2024)
SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment Optimization
by: Li, Wanhua, et al.
Published: (2024)
by: Li, Wanhua, et al.
Published: (2024)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024)
by: Ma, Chenyang, et al.
Published: (2024)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
by: Peng, Qihang, et al.
Published: (2025)
by: Peng, Qihang, et al.
Published: (2025)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
by: Jia, Yueru, et al.
Published: (2024)
by: Jia, Yueru, et al.
Published: (2024)
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
by: Yuan, Zhihao, et al.
Published: (2025)
by: Yuan, Zhihao, et al.
Published: (2025)
No Time to Train: Empowering Non-Parametric Networks for Few-shot 3D Scene Segmentation
by: Zhu, Xiangyang, et al.
Published: (2024)
by: Zhu, Xiangyang, et al.
Published: (2024)
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
by: Wei, Xiaobao, et al.
Published: (2023)
by: Wei, Xiaobao, et al.
Published: (2023)
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory
by: Cao, Jianbao, et al.
Published: (2026)
by: Cao, Jianbao, et al.
Published: (2026)
SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation
by: Li, Yuan, et al.
Published: (2026)
by: Li, Yuan, et al.
Published: (2026)
Training-free Regional Prompting for Diffusion Transformers
by: Chen, Anthony, et al.
Published: (2024)
by: Chen, Anthony, et al.
Published: (2024)
Prompt fidelity of ChatGPT4o / Dall-E3 text-to-image visualisations
by: Spennemann, Dirk HR
Published: (2025)
by: Spennemann, Dirk HR
Published: (2025)
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
by: Jin, Zhao, et al.
Published: (2025)
by: Jin, Zhao, et al.
Published: (2025)
PF3Det: A Prompted Foundation Feature Assisted Visual LiDAR 3D Detector
by: Li, Kaidong, et al.
Published: (2025)
by: Li, Kaidong, et al.
Published: (2025)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations
by: Yuan, Jiangye, et al.
Published: (2026)
by: Yuan, Jiangye, et al.
Published: (2026)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
Similar Items
-
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
by: Liu, Dingning, et al.
Published: (2024) -
Learning Geometry-Guided Depth via Projective Modeling for Monocular 3D Object Detection
by: Zhang, Yinmin, et al.
Published: (2021) -
Referring Video Object Segmentation with Cross-Modality Proxy Queries
by: Sun, Baoli, et al.
Published: (2025) -
Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth Completion
by: Chen, Shenglun, et al.
Published: (2025) -
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
by: Sun, Baoli, et al.
Published: (2025)