POMA-3D: The Point Map Way to 3D Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Ye, Luo, Weixun, Huang, Ranran, Jing, Junpeng, Mikolajczyk, Krystian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2026)
by: Mao, Ye, et al.
Published: (2026)
From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
by: Huang, Ranran, et al.
Published: (2026)
by: Huang, Ranran, et al.
Published: (2026)
Hypo3D: Exploring Hypothetical Reasoning in 3D
by: Mao, Ye, et al.
Published: (2025)
by: Mao, Ye, et al.
Published: (2025)
Stereo Any Video: Temporally Consistent Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)
by: Jing, Junpeng, et al.
Published: (2025)
Lite Any Stereo: Efficient Zero-Shot Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)
by: Jing, Junpeng, et al.
Published: (2025)
OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images
by: Mao, Ye, et al.
Published: (2024)
by: Mao, Ye, et al.
Published: (2024)
No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
by: Huang, Ranran, et al.
Published: (2025)
by: Huang, Ranran, et al.
Published: (2025)
SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
by: Huang, Ranran, et al.
Published: (2025)
by: Huang, Ranran, et al.
Published: (2025)
Match-Stereo-Videos: Bidirectional Alignment for Consistent Dynamic Stereo Matching
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
Match Stereo Videos via Bidirectional Alignment
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
Understanding the Role of the Projector in Knowledge Distillation
by: Miles, Roy, et al.
Published: (2023)
by: Miles, Roy, et al.
Published: (2023)
UCorr: Wire Detection and Depth Estimation for Autonomous Drones
by: Kolbeinsson, Benedikt, et al.
Published: (2025)
by: Kolbeinsson, Benedikt, et al.
Published: (2025)
Multi-Class Segmentation from Aerial Views using Recursive Noise Diffusion
by: Kolbeinsson, Benedikt, et al.
Published: (2022)
by: Kolbeinsson, Benedikt, et al.
Published: (2022)
DDOS: The Drone Depth and Obstacle Segmentation Dataset
by: Kolbeinsson, Benedikt, et al.
Published: (2023)
by: Kolbeinsson, Benedikt, et al.
Published: (2023)
Language-Based Depth Hints for Monocular Depth Estimation
by: Auty, Dylan, et al.
Published: (2024)
by: Auty, Dylan, et al.
Published: (2024)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
SDesc3D: Towards Layout-Aware 3D Indoor Scene Generation from Short Descriptions
by: Feng, Jie, et al.
Published: (2026)
by: Feng, Jie, et al.
Published: (2026)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
Is Your LiDAR Placement Optimized for 3D Scene Understanding?
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
Scene Reconstruction as Mapping Priors for 3D Detection
by: Fu, Yang, et al.
Published: (2026)
by: Fu, Yang, et al.
Published: (2026)
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
by: Huang, Wencan, et al.
Published: (2025)
by: Huang, Wencan, et al.
Published: (2025)
AVS-Net: Point Sampling with Adaptive Voxel Size for 3D Scene Understanding
by: Yang, Hongcheng, et al.
Published: (2024)
by: Yang, Hongcheng, et al.
Published: (2024)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
by: Huang, Junsheng, et al.
Published: (2025)
by: Huang, Junsheng, et al.
Published: (2025)
SAM-Guided Masked Token Prediction for 3D Scene Understanding
by: Chen, Zhimin, et al.
Published: (2024)
by: Chen, Zhimin, et al.
Published: (2024)
QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
by: Mehan, Yash, et al.
Published: (2024)
by: Mehan, Yash, et al.
Published: (2024)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
by: Halacheva, Anna-Maria, et al.
Published: (2024)
by: Halacheva, Anna-Maria, et al.
Published: (2024)
Open-Vocabulary Octree-Graph for 3D Scene Understanding
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
SceneGPT: A Language Model for 3D Scene Understanding
by: Chandhok, Shivam
Published: (2024)
by: Chandhok, Shivam
Published: (2024)
R3DS: Reality-linked 3D Scenes for Panoramic Scene Understanding
by: Wu, Qirui, et al.
Published: (2024)
by: Wu, Qirui, et al.
Published: (2024)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)
by: Zheng, Duo, et al.
Published: (2024)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
by: Linghu, Xiongkun, et al.
Published: (2026)
by: Linghu, Xiongkun, et al.
Published: (2026)
Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
Unified Semantic Transformer for 3D Scene Understanding
by: Koch, Sebastian, et al.
Published: (2025)
by: Koch, Sebastian, et al.
Published: (2025)
A Unified Framework for 3D Scene Understanding
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
3D Question Answering for City Scene Understanding
by: Sun, Penglei, et al.
Published: (2024)
by: Sun, Penglei, et al.
Published: (2024)
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
Similar Items
-
Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2026) -
From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
by: Huang, Ranran, et al.
Published: (2026) -
Hypo3D: Exploring Hypothetical Reasoning in 3D
by: Mao, Ye, et al.
Published: (2025) -
Stereo Any Video: Temporally Consistent Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025) -
Lite Any Stereo: Efficient Zero-Shot Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)