ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Junliang, Wang, Zhengyi, Zhao, Ruowen, Xie, Shenghao, Zhu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
by: Ye, Junliang, et al.
Published: (2025)
by: Ye, Junliang, et al.
Published: (2025)
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
by: Qi, Zekun, et al.
Published: (2024)
by: Qi, Zekun, et al.
Published: (2024)
FlexiDreamer: Single Image-to-3D Generation with FlexiCubes
by: Zhao, Ruowen, et al.
Published: (2024)
by: Zhao, Ruowen, et al.
Published: (2024)
DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning
by: Zhao, Ruowen, et al.
Published: (2025)
by: Zhao, Ruowen, et al.
Published: (2025)
Compositional 3D-aware Video Generation with LLM Director
by: Zhu, Hanxin, et al.
Published: (2024)
by: Zhu, Hanxin, et al.
Published: (2024)
DreamReward: Text-to-3D Generation with Human Preference
by: Ye, Junliang, et al.
Published: (2024)
by: Ye, Junliang, et al.
Published: (2024)
In-Place Panoptic Radiance Field Segmentation with Perceptual Prior for 3D Scene Understanding
by: Li, Shenghao
Published: (2024)
by: Li, Shenghao
Published: (2024)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
by: Ye, Hanrong, et al.
Published: (2025)
by: Ye, Hanrong, et al.
Published: (2025)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
by: Zong, Zhuofan, et al.
Published: (2024)
by: Zong, Zhuofan, et al.
Published: (2024)
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
ArtLLM: Generating Articulated Assets via 3D LLM
by: Wang, Penghao, et al.
Published: (2026)
by: Wang, Penghao, et al.
Published: (2026)
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
by: Zhao, Shifang, et al.
Published: (2025)
by: Zhao, Shifang, et al.
Published: (2025)
Part-X-MLLM: Part-aware 3D Multimodal Large Language Model
by: Wang, Chunshi, et al.
Published: (2025)
by: Wang, Chunshi, et al.
Published: (2025)
Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
by: Zhou, Zhiwang, et al.
Published: (2025)
by: Zhou, Zhiwang, et al.
Published: (2025)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
by: Li, Lijiang, et al.
Published: (2026)
by: Li, Lijiang, et al.
Published: (2026)
EmoLLM: Multimodal Emotional Understanding Meets Large Language Models
by: Yang, Qu, et al.
Published: (2024)
by: Yang, Qu, et al.
Published: (2024)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning
by: Qiu, Liuxiang, et al.
Published: (2026)
by: Qiu, Liuxiang, et al.
Published: (2026)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
by: Ye, Chongjie, et al.
Published: (2026)
by: Ye, Chongjie, et al.
Published: (2026)
BrepLLM: Native Boundary Representation Understanding with Large Language Models
by: Deng, Liyuan, et al.
Published: (2025)
by: Deng, Liyuan, et al.
Published: (2025)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
A Survey on Image Quality Assessment: Insights, Analysis, and Future Outlook
by: Ma, Chengqian, et al.
Published: (2025)
by: Ma, Chengqian, et al.
Published: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
by: Kang, Hengrui, et al.
Published: (2025)
by: Kang, Hengrui, et al.
Published: (2025)
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
Towards Native Generative Model for 3D Head Avatar
by: Zhuang, Yiyu, et al.
Published: (2024)
by: Zhuang, Yiyu, et al.
Published: (2024)
Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels
by: Wang, Yikai, et al.
Published: (2024)
by: Wang, Yikai, et al.
Published: (2024)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers
by: Yang, Zongyuan, et al.
Published: (2026)
by: Yang, Zongyuan, et al.
Published: (2026)
V3D: Video Diffusion Models are Effective 3D Generators
by: Chen, Zilong, et al.
Published: (2024)
by: Chen, Zilong, et al.
Published: (2024)
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation
by: Lee, Suhyeon, et al.
Published: (2023)
by: Lee, Suhyeon, et al.
Published: (2023)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
MicroDreamer: Efficient 3D Generation in $\sim$20 Seconds by Score-based Iterative Reconstruction
by: Chen, Luxi, et al.
Published: (2024)
by: Chen, Luxi, et al.
Published: (2024)
ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
by: Tian, Xueyun, et al.
Published: (2026)
by: Tian, Xueyun, et al.
Published: (2026)
Similar Items
-
NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
by: Ye, Junliang, et al.
Published: (2025) -
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
by: Qi, Zekun, et al.
Published: (2024) -
FlexiDreamer: Single Image-to-3D Generation with FlexiCubes
by: Zhao, Ruowen, et al.
Published: (2024) -
DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning
by: Zhao, Ruowen, et al.
Published: (2025) -
Compositional 3D-aware Video Generation with LLM Director
by: Zhu, Hanxin, et al.
Published: (2024)