3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Xiong, Haomiao, Zhuge, Yunzhi, Zhu, Jiawen, Zhang, Lu, Lu, Huchuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
by: Zhang, Lu, et al.
Published: (2025)
by: Zhang, Lu, et al.
Published: (2025)
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
by: Feng, Xiyan, et al.
Published: (2026)
by: Feng, Xiyan, et al.
Published: (2026)
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
by: Zhuge, Yunzhi, et al.
Published: (2025)
by: Zhuge, Yunzhi, et al.
Published: (2025)
Parameter Aware Mamba Model for Multi-task Dense Prediction
by: Yu, Xinzhuo, et al.
Published: (2025)
by: Yu, Xinzhuo, et al.
Published: (2025)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
by: Wu, Jiannan, et al.
Published: (2024)
by: Wu, Jiannan, et al.
Published: (2024)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
by: Liu, Qing'an, et al.
Published: (2026)
by: Liu, Qing'an, et al.
Published: (2026)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Learning Universal Features for Generalizable Image Forgery Localization
by: Zhao, Hengrun, et al.
Published: (2025)
by: Zhao, Hengrun, et al.
Published: (2025)
Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting
by: Zhu, Runsong, et al.
Published: (2025)
by: Zhu, Runsong, et al.
Published: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
StableIdentity: Inserting Anybody into Anywhere at First Sight
by: Wang, Qinghe, et al.
Published: (2024)
by: Wang, Qinghe, et al.
Published: (2024)
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models
by: Thomas, Hugues, et al.
Published: (2025)
by: Thomas, Hugues, et al.
Published: (2025)
Regularizing Subspace Redundancy of Low-Rank Adaptation
by: Zhu, Yue, et al.
Published: (2025)
by: Zhu, Yue, et al.
Published: (2025)
VLM-3D:End-to-End Vision-Language Models for Open-World 3D Perception
by: Chang, Fuhao, et al.
Published: (2025)
by: Chang, Fuhao, et al.
Published: (2025)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting
by: Yang, Yicheng, et al.
Published: (2024)
by: Yang, Yicheng, et al.
Published: (2024)
End-to-End 3D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
by: Yang, Zhenwei, et al.
Published: (2025)
by: Yang, Zhenwei, et al.
Published: (2025)
Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding
by: Zemskova, Tatiana, et al.
Published: (2024)
by: Zemskova, Tatiana, et al.
Published: (2024)
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
by: Fu, Rao, et al.
Published: (2024)
by: Fu, Rao, et al.
Published: (2024)
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
Multimodal 3D Reasoning Segmentation with Complex Scenes
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
by: You, Ling, et al.
Published: (2025)
by: You, Ling, et al.
Published: (2025)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)
by: Wang, Linhan, et al.
Published: (2026)
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
by: Song, Nan, et al.
Published: (2025)
by: Song, Nan, et al.
Published: (2025)
SceneGPT: A Language Model for 3D Scene Understanding
by: Chandhok, Shivam
Published: (2024)
by: Chandhok, Shivam
Published: (2024)
SceneLCM: End-to-End Layout-Guided Interactive Indoor Scene Generation with Latent Consistency Model
by: Lin, Yangkai, et al.
Published: (2025)
by: Lin, Yangkai, et al.
Published: (2025)
Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation
by: Lei, Biwen, et al.
Published: (2025)
by: Lei, Biwen, et al.
Published: (2025)
End-to-End Rate-Distortion Optimized 3D Gaussian Representation
by: Wang, Henan, et al.
Published: (2024)
by: Wang, Henan, et al.
Published: (2024)
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
Similar Items
-
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
by: Xiong, Haomiao, et al.
Published: (2025) -
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
by: Zhang, Wenbo, et al.
Published: (2024) -
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
by: Zhang, Lu, et al.
Published: (2025) -
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
by: Feng, Xiyan, et al.
Published: (2026) -
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)