Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Wencan, Liu, Daizong, Hu, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
di: Liu, Daizong, et al.
Pubblicazione: (2024)
di: Liu, Daizong, et al.
Pubblicazione: (2024)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation
di: Cheng, Wencan, et al.
Pubblicazione: (2026)
di: Cheng, Wencan, et al.
Pubblicazione: (2026)
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
di: Li, Zeju, et al.
Pubblicazione: (2024)
di: Li, Zeju, et al.
Pubblicazione: (2024)
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
di: Du, Bi'an, et al.
Pubblicazione: (2026)
di: Du, Bi'an, et al.
Pubblicazione: (2026)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
di: Yu, Hanxun, et al.
Pubblicazione: (2025)
di: Yu, Hanxun, et al.
Pubblicazione: (2025)
Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
Fast SceneScript: Fast and Accurate Language-Based 3D Scene Understanding via Multi-Token Prediction
di: Yin, Ruihong, et al.
Pubblicazione: (2025)
di: Yin, Ruihong, et al.
Pubblicazione: (2025)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
di: Li, Chen, et al.
Pubblicazione: (2025)
di: Li, Chen, et al.
Pubblicazione: (2025)
Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
di: Xue, Jintang, et al.
Pubblicazione: (2025)
di: Xue, Jintang, et al.
Pubblicazione: (2025)
Multi-modal Situated Reasoning in 3D Scenes
di: Linghu, Xiongkun, et al.
Pubblicazione: (2024)
di: Linghu, Xiongkun, et al.
Pubblicazione: (2024)
SceneGPT: A Language Model for 3D Scene Understanding
di: Chandhok, Shivam
Pubblicazione: (2024)
di: Chandhok, Shivam
Pubblicazione: (2024)
Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models
di: Thomas, Hugues, et al.
Pubblicazione: (2025)
di: Thomas, Hugues, et al.
Pubblicazione: (2025)
Swin3D++: Effective Multi-Source Pretraining for 3D Indoor Scene Understanding
di: Yang, Yu-Qi, et al.
Pubblicazione: (2024)
di: Yang, Yu-Qi, et al.
Pubblicazione: (2024)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
di: Huang, Haifeng, et al.
Pubblicazione: (2023)
di: Huang, Haifeng, et al.
Pubblicazione: (2023)
POMA-3D: The Point Map Way to 3D Scene Understanding
di: Mao, Ye, et al.
Pubblicazione: (2025)
di: Mao, Ye, et al.
Pubblicazione: (2025)
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
di: Ji, Jiayi, et al.
Pubblicazione: (2023)
di: Ji, Jiayi, et al.
Pubblicazione: (2023)
SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models
di: Zhang, Yue, et al.
Pubblicazione: (2024)
di: Zhang, Yue, et al.
Pubblicazione: (2024)
Fast and Efficient: Mask Neural Fields for 3D Scene Segmentation
di: Gao, Zihan, et al.
Pubblicazione: (2024)
di: Gao, Zihan, et al.
Pubblicazione: (2024)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
di: Xiong, Haomiao, et al.
Pubblicazione: (2025)
di: Xiong, Haomiao, et al.
Pubblicazione: (2025)
3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding
di: Zemskova, Tatiana, et al.
Pubblicazione: (2024)
di: Zemskova, Tatiana, et al.
Pubblicazione: (2024)
Language-Assisted 3D Scene Understanding
di: Wu, Yanmin, et al.
Pubblicazione: (2023)
di: Wu, Yanmin, et al.
Pubblicazione: (2023)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
di: Huang, Junming, et al.
Pubblicazione: (2026)
di: Huang, Junming, et al.
Pubblicazione: (2026)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
Unified Scene Representation and Reconstruction for 3D Large Language Models
di: Chu, Tao, et al.
Pubblicazione: (2024)
di: Chu, Tao, et al.
Pubblicazione: (2024)
OpenSU3D: Open World 3D Scene Understanding using Foundation Models
di: Mohiuddin, Rafay, et al.
Pubblicazione: (2024)
di: Mohiuddin, Rafay, et al.
Pubblicazione: (2024)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding
di: Zhu, Wenxuan, et al.
Pubblicazione: (2025)
di: Zhu, Wenxuan, et al.
Pubblicazione: (2025)
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
di: Yuan, Zhihao, et al.
Pubblicazione: (2025)
di: Yuan, Zhihao, et al.
Pubblicazione: (2025)
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding
di: Zheng, Henry, et al.
Pubblicazione: (2026)
di: Zheng, Henry, et al.
Pubblicazione: (2026)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
di: Kong, Lingdong, et al.
Pubblicazione: (2024)
di: Kong, Lingdong, et al.
Pubblicazione: (2024)
IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation
di: Zhou, Wenxu, et al.
Pubblicazione: (2025)
di: Zhou, Wenxu, et al.
Pubblicazione: (2025)
LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds
di: Bang, Jaehun, et al.
Pubblicazione: (2026)
di: Bang, Jaehun, et al.
Pubblicazione: (2026)
SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment
di: Xu, Qi, et al.
Pubblicazione: (2025)
di: Xu, Qi, et al.
Pubblicazione: (2025)
A Unified Framework for 3D Scene Understanding
di: Xu, Wei, et al.
Pubblicazione: (2024)
di: Xu, Wei, et al.
Pubblicazione: (2024)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
di: Halacheva, Anna-Maria, et al.
Pubblicazione: (2024)
di: Halacheva, Anna-Maria, et al.
Pubblicazione: (2024)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
di: Qi, Zhangyang, et al.
Pubblicazione: (2025)
di: Qi, Zhangyang, et al.
Pubblicazione: (2025)
Hard-Label Black-Box Attacks on 3D Point Clouds
di: Liu, Daizong, et al.
Pubblicazione: (2024)
di: Liu, Daizong, et al.
Pubblicazione: (2024)
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
di: Tai, Hanchen, et al.
Pubblicazione: (2024)
di: Tai, Hanchen, et al.
Pubblicazione: (2024)
NIS-SLAM: Neural Implicit Semantic RGB-D SLAM for 3D Consistent Scene Understanding
di: Zhai, Hongjia, et al.
Pubblicazione: (2024)
di: Zhai, Hongjia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
di: Liu, Daizong, et al.
Pubblicazione: (2024) -
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
di: Liu, Yang, et al.
Pubblicazione: (2024) -
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation
di: Cheng, Wencan, et al.
Pubblicazione: (2026) -
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
di: Li, Zeju, et al.
Pubblicazione: (2024) -
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
di: Du, Bi'an, et al.
Pubblicazione: (2026)