Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Hanxun, Li, Wentong, Wang, Song, Chen, Junbo, Zhu, Jianke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
von: Liu, Hongyuan, et al.
Veröffentlicht: (2025)
von: Liu, Hongyuan, et al.
Veröffentlicht: (2025)
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
von: Li, Zeju, et al.
Veröffentlicht: (2024)
von: Li, Zeju, et al.
Veröffentlicht: (2024)
Not All Voxels Are Equal: Hardness-Aware Semantic Scene Completion with Self-Distillation
von: Wang, Song, et al.
Veröffentlicht: (2024)
von: Wang, Song, et al.
Veröffentlicht: (2024)
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
Label-efficient Semantic Scene Completion with Scribble Annotations
von: Wang, Song, et al.
Veröffentlicht: (2024)
von: Wang, Song, et al.
Veröffentlicht: (2024)
Osprey: Pixel Understanding with Visual Instruction Tuning
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
Uncertainty-Instructed Structure Injection for Generalizable HD Map Construction
von: Liu, Xiaolu, et al.
Veröffentlicht: (2025)
von: Liu, Xiaolu, et al.
Veröffentlicht: (2025)
MGMap: Mask-Guided Learning for Online Vectorized HD Map Construction
von: Liu, Xiaolu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaolu, et al.
Veröffentlicht: (2024)
Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding
von: Zheng, Hongpei, et al.
Veröffentlicht: (2025)
von: Zheng, Hongpei, et al.
Veröffentlicht: (2025)
ReliOcc: Towards Reliable Semantic Occupancy Prediction via Uncertainty Learning
von: Wang, Song, et al.
Veröffentlicht: (2024)
von: Wang, Song, et al.
Veröffentlicht: (2024)
IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection
von: Yin, Junbo, et al.
Veröffentlicht: (2024)
von: Yin, Junbo, et al.
Veröffentlicht: (2024)
Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion
von: Liu, Enyu, et al.
Veröffentlicht: (2025)
von: Liu, Enyu, et al.
Veröffentlicht: (2025)
HO-Gaussian: Hybrid Optimization of 3D Gaussian Splatting for Urban Scenes
von: Li, Zhuopeng, et al.
Veröffentlicht: (2024)
von: Li, Zhuopeng, et al.
Veröffentlicht: (2024)
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2026)
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2026)
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
von: Huang, Wencan, et al.
Veröffentlicht: (2025)
von: Huang, Wencan, et al.
Veröffentlicht: (2025)
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
von: Wang, Xiaoye, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoye, et al.
Veröffentlicht: (2025)
Unlocking Dense Metric Depth Estimation in VLMs
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
Fine-Grained Multi-View Hand Reconstruction Using Inverse Rendering
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
SAI3D: Segment Any Instance in 3D Scenes
von: Yin, Yingda, et al.
Veröffentlicht: (2023)
von: Yin, Yingda, et al.
Veröffentlicht: (2023)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
von: Lee, Yonghan, et al.
Veröffentlicht: (2026)
von: Lee, Yonghan, et al.
Veröffentlicht: (2026)
UnScene3D: Unsupervised 3D Instance Segmentation for Indoor Scenes
von: Rozenberszki, David, et al.
Veröffentlicht: (2023)
von: Rozenberszki, David, et al.
Veröffentlicht: (2023)
MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
von: Yang, Bin, et al.
Veröffentlicht: (2026)
von: Yang, Bin, et al.
Veröffentlicht: (2026)
Swin3D++: Effective Multi-Source Pretraining for 3D Indoor Scene Understanding
von: Yang, Yu-Qi, et al.
Veröffentlicht: (2024)
von: Yang, Yu-Qi, et al.
Veröffentlicht: (2024)
Q-Adapt: Adapting LMM for Visual Quality Assessment with Progressive Instruction Tuning
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
AutoInst: Automatic Instance-Based Segmentation of LiDAR 3D Scans
von: Perauer, Cedric, et al.
Veröffentlicht: (2024)
von: Perauer, Cedric, et al.
Veröffentlicht: (2024)
HVOFusion: Incremental Mesh Reconstruction Using Hybrid Voxel Octree
von: Liu, Shaofan, et al.
Veröffentlicht: (2024)
von: Liu, Shaofan, et al.
Veröffentlicht: (2024)
3D Question Answering for City Scene Understanding
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
AG$^2$aussian: Anchor-Graph Structured Gaussian Splatting for Instance-Level 3D Scene Understanding and Editing
von: Wang, Zhaonan, et al.
Veröffentlicht: (2025)
von: Wang, Zhaonan, et al.
Veröffentlicht: (2025)
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
von: Ji, Jiayi, et al.
Veröffentlicht: (2023)
von: Ji, Jiayi, et al.
Veröffentlicht: (2023)
UDA4Inst: Unsupervised Domain Adaptation for Instance Segmentation
von: Guo, Yachan, et al.
Veröffentlicht: (2024)
von: Guo, Yachan, et al.
Veröffentlicht: (2024)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
von: Zhang, Yisu, et al.
Veröffentlicht: (2026)
von: Zhang, Yisu, et al.
Veröffentlicht: (2026)
Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing
von: Liu, Xiaolu, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolu, et al.
Veröffentlicht: (2026)
Multi-modal Situated Reasoning in 3D Scenes
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
von: Li, Zechuan, et al.
Veröffentlicht: (2025)
von: Li, Zechuan, et al.
Veröffentlicht: (2025)
Instance Tracking in 3D Scenes from Egocentric Videos
von: Zhao, Yunhan, et al.
Veröffentlicht: (2023)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2023)
ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
von: Yu, Hanxun, et al.
Veröffentlicht: (2026) -
InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
von: Liu, Hongyuan, et al.
Veröffentlicht: (2025) -
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
von: Li, Zeju, et al.
Veröffentlicht: (2024) -
Not All Voxels Are Equal: Hardness-Aware Semantic Scene Completion with Self-Distillation
von: Wang, Song, et al.
Veröffentlicht: (2024) -
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
von: Shi, Zhan, et al.
Veröffentlicht: (2025)