UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
Fuente:
arXiv
Saved in:
| Main Authors: | He, Qingdong, Peng, Jinlong, Jiang, Zhengkai, Wu, Kai, Ji, Xiaozhong, Zhang, Jiangning, Wang, Yabiao, Wang, Chengjie, Chen, Mingang, Wu, Yunsheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
by: Tai, Hanchen, et al.
Published: (2024)
by: Tai, Hanchen, et al.
Published: (2024)
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
PointSeg: A Training-Free Paradigm for 3D Scene Segmentation via Foundation Models
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
UniM$^2$AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
by: Zou, Jian, et al.
Published: (2023)
by: Zou, Jian, et al.
Published: (2023)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection
by: Wang, Haoxuan, et al.
Published: (2024)
by: Wang, Haoxuan, et al.
Published: (2024)
PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
M3DM-NR: RGB-3D Noisy-Resistant Industrial Anomaly Detection via Multimodal Denoising
by: Wang, Chengjie, et al.
Published: (2024)
by: Wang, Chengjie, et al.
Published: (2024)
AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval
by: Zhang, Sihe, et al.
Published: (2024)
by: Zhang, Sihe, et al.
Published: (2024)
Leveraging Fine-Grained Information and Noise Decoupling for Remote Sensing Change Detection
by: Du, Qiangang, et al.
Published: (2024)
by: Du, Qiangang, et al.
Published: (2024)
UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer
by: Wang, Haoxuan, et al.
Published: (2025)
by: Wang, Haoxuan, et al.
Published: (2025)
Open-Vocabulary Octree-Graph for 3D Scene Understanding
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024)
by: Fu, Chencan, et al.
Published: (2024)
Self-supervised Feature Adaptation for 3D Industrial Anomaly Detection
by: Tu, Yuanpeng, et al.
Published: (2024)
by: Tu, Yuanpeng, et al.
Published: (2024)
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
UniPLV: Towards Label-Efficient Open-World 3D Scene Understanding by Regional Visual Language Supervision
by: Wang, Yuru, et al.
Published: (2024)
by: Wang, Yuru, et al.
Published: (2024)
PiT: Progressive Diffusion Transformer
by: Wu, Jiafu, et al.
Published: (2025)
by: Wu, Jiafu, et al.
Published: (2025)
OGScene3D: Incremental Open-Vocabulary 3D Gaussian Scene Graph Mapping for Scene Understanding
by: Zhu, Siting, et al.
Published: (2026)
by: Zhu, Siting, et al.
Published: (2026)
OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding
by: Engelmann, Francis, et al.
Published: (2024)
by: Engelmann, Francis, et al.
Published: (2024)
MARRS: Masked Autoregressive Unit-based Reaction Synthesis
by: Wang, Yabiao, et al.
Published: (2025)
by: Wang, Yabiao, et al.
Published: (2025)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection
by: Chow, Adrian, et al.
Published: (2025)
by: Chow, Adrian, et al.
Published: (2025)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections
by: Wang, Fangjinhua, et al.
Published: (2023)
by: Wang, Fangjinhua, et al.
Published: (2023)
FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on
by: Jiang, Boyuan, et al.
Published: (2024)
by: Jiang, Boyuan, et al.
Published: (2024)
Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
by: Wang, Pengfei, et al.
Published: (2024)
by: Wang, Pengfei, et al.
Published: (2024)
UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation
by: Jin, Bu, et al.
Published: (2025)
by: Jin, Bu, et al.
Published: (2025)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
by: Li, Ruihuang, et al.
Published: (2024)
by: Li, Ruihuang, et al.
Published: (2024)
Collaborative Dynamic 3D Scene Graphs for Open-Vocabulary Urban Scene Understanding
by: Steinke, Tim, et al.
Published: (2025)
by: Steinke, Tim, et al.
Published: (2025)
ImOV3D: Learning Open-Vocabulary Point Clouds 3D Object Detection from Only 2D Images
by: Yang, Timing, et al.
Published: (2024)
by: Yang, Timing, et al.
Published: (2024)
OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations
by: Kassab, Christina, et al.
Published: (2025)
by: Kassab, Christina, et al.
Published: (2025)
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
by: Lee, Seungjun, et al.
Published: (2026)
by: Lee, Seungjun, et al.
Published: (2026)
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
by: Huang, Sheng-Yu, et al.
Published: (2026)
by: Huang, Sheng-Yu, et al.
Published: (2026)
UniMesh: Unifying 3D Mesh Understanding and Generation
by: Huang, Peng, et al.
Published: (2026)
by: Huang, Peng, et al.
Published: (2026)
DiffuMatting: Synthesizing Arbitrary Objects with Matting-level Annotation
by: Hu, Xiaobin, et al.
Published: (2024)
by: Hu, Xiaobin, et al.
Published: (2024)
Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection
by: He, Liren, et al.
Published: (2024)
by: He, Liren, et al.
Published: (2024)
Similar Items
-
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
by: Tai, Hanchen, et al.
Published: (2024) -
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
by: Wang, Zhenyu, et al.
Published: (2024) -
PointSeg: A Training-Free Paradigm for 3D Scene Segmentation via Foundation Models
by: He, Qingdong, et al.
Published: (2024) -
UniM$^2$AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
by: Zou, Jian, et al.
Published: (2023) -
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)