Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junhao, Li, Xiang, Ye, Xiaojun, Li, Chao, Fan, Zhaoxin, Zhao, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ultraman: Single Image 3D Human Reconstruction with Ultra Speed and Detail
by: Chen, Mingjin, et al.
Published: (2024)
by: Chen, Mingjin, et al.
Published: (2024)
F-LMM: Grounding Frozen Large Multimodal Models
by: Wu, Size, et al.
Published: (2024)
by: Wu, Size, et al.
Published: (2024)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
by: Li, Feng, et al.
Published: (2024)
by: Li, Feng, et al.
Published: (2024)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025)
by: Yu, Hanxun, et al.
Published: (2025)
LMM-Det: Make Large Multimodal Models Excel in Object Detection
by: Li, Jincheng, et al.
Published: (2025)
by: Li, Jincheng, et al.
Published: (2025)
LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
DuoGen: Towards General Purpose Interleaved Multimodal Generation
by: Shi, Min, et al.
Published: (2026)
by: Shi, Min, et al.
Published: (2026)
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
by: Shukla, Shreya, et al.
Published: (2025)
by: Shukla, Shreya, et al.
Published: (2025)
Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
by: Li, Bao, et al.
Published: (2025)
by: Li, Bao, et al.
Published: (2025)
Improving Multimodal Hateful Meme Detection Exploiting LMM-Generated Knowledge
by: Tzelepi, Maria, et al.
Published: (2025)
by: Tzelepi, Maria, et al.
Published: (2025)
HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
by: Chen, Mingjin, et al.
Published: (2026)
by: Chen, Mingjin, et al.
Published: (2026)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
by: Ge, Qihang, et al.
Published: (2024)
by: Ge, Qihang, et al.
Published: (2024)
LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue
by: Li, Chaoyue, et al.
Published: (2026)
by: Li, Chaoyue, et al.
Published: (2026)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
by: Liao, Chao, et al.
Published: (2025)
by: Liao, Chao, et al.
Published: (2025)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
by: He, Bo, et al.
Published: (2024)
by: He, Bo, et al.
Published: (2024)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
MLPHand: Real Time Multi-View 3D Hand Mesh Reconstruction via MLP Modeling
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
by: Li, Wenbin, et al.
Published: (2026)
by: Li, Wenbin, et al.
Published: (2026)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
Towards Geometric and Textural Consistency 3D Scene Generation via Single Image-guided Model Generation and Layout Optimization
by: Tang, Xiang, et al.
Published: (2025)
by: Tang, Xiang, et al.
Published: (2025)
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
by: Liu, Qingyang, et al.
Published: (2026)
by: Liu, Qingyang, et al.
Published: (2026)
Recent Advances in 3D Object and Scene Generation: A Survey
by: Tang, Xiang, et al.
Published: (2025)
by: Tang, Xiang, et al.
Published: (2025)
LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
by: Wang, Shengyuan, et al.
Published: (2025)
by: Wang, Shengyuan, et al.
Published: (2025)
InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling
by: Javed, Muhammad Gohar, et al.
Published: (2024)
by: Javed, Muhammad Gohar, et al.
Published: (2024)
InScope: A New Real-world 3D Infrastructure-side Collaborative Perception Dataset for Open Traffic Scenarios
by: Zhang, Xiaofei, et al.
Published: (2024)
by: Zhang, Xiaofei, et al.
Published: (2024)
SiMO: Single-Modality-Operable Multimodal Collaborative Perception
by: Wen, Jiageng, et al.
Published: (2026)
by: Wen, Jiageng, et al.
Published: (2026)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
by: Yin, Shaofeng, et al.
Published: (2026)
by: Yin, Shaofeng, et al.
Published: (2026)
Loom: Diffusion-Transformer for Interleaved Generation
by: Ye, Mingcheng, et al.
Published: (2025)
by: Ye, Mingcheng, et al.
Published: (2025)
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
by: Nie, Ming, et al.
Published: (2026)
by: Nie, Ming, et al.
Published: (2026)
ForgeDreamer: Industrial Text-to-3D Generation with Multi-Expert LoRA and Cross-View Hypergraph
by: Cai, Junhao, et al.
Published: (2026)
by: Cai, Junhao, et al.
Published: (2026)
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
by: Gu, Jiawei, et al.
Published: (2025)
by: Gu, Jiawei, et al.
Published: (2025)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
by: Li, Chenxu, et al.
Published: (2025)
by: Li, Chenxu, et al.
Published: (2025)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
by: Wu, Shengqiong, et al.
Published: (2026)
by: Wu, Shengqiong, et al.
Published: (2026)
Enhancing Weakly Supervised 3D Medical Image Segmentation through Probabilistic-aware Learning
by: Jiang, Runmin, et al.
Published: (2024)
by: Jiang, Runmin, et al.
Published: (2024)
Similar Items
-
Ultraman: Single Image 3D Human Reconstruction with Ultra Speed and Detail
by: Chen, Mingjin, et al.
Published: (2024) -
F-LMM: Grounding Frozen Large Multimodal Models
by: Wu, Size, et al.
Published: (2024) -
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
by: Li, Feng, et al.
Published: (2024) -
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025) -
LMM-Det: Make Large Multimodal Models Excel in Object Detection
by: Li, Jincheng, et al.
Published: (2025)