Collaborative Multi-Modal Coding for High-Quality 3D Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Ziang, Chen, Zhaoxi, Pan, Liang, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PhysX-3D: Physical-Grounded 3D Asset Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
von: Cao, Ziang, et al.
Veröffentlicht: (2026)
von: Cao, Ziang, et al.
Veröffentlicht: (2026)
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2024)
von: Cao, Ziang, et al.
Veröffentlicht: (2024)
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
von: Liu, Zixian, et al.
Veröffentlicht: (2026)
von: Liu, Zixian, et al.
Veröffentlicht: (2026)
3DTopia: Large Text-to-3D Generation Model with Hybrid Diffusion Priors
von: Hong, Fangzhou, et al.
Veröffentlicht: (2024)
von: Hong, Fangzhou, et al.
Veröffentlicht: (2024)
Generative Gaussian Splatting for Unbounded 3D City Generation
von: Xie, Haozhe, et al.
Veröffentlicht: (2024)
von: Xie, Haozhe, et al.
Veröffentlicht: (2024)
3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion
von: Chen, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Chen, Zhaoxi, et al.
Veröffentlicht: (2024)
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
von: Xie, Haozhe, et al.
Veröffentlicht: (2023)
von: Xie, Haozhe, et al.
Veröffentlicht: (2023)
FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls
von: Hu, Tao, et al.
Veröffentlicht: (2024)
von: Hu, Tao, et al.
Veröffentlicht: (2024)
LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation
von: Tang, Jiaxiang, et al.
Veröffentlicht: (2024)
von: Tang, Jiaxiang, et al.
Veröffentlicht: (2024)
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
von: Lv, Zhengyao, et al.
Veröffentlicht: (2025)
von: Lv, Zhengyao, et al.
Veröffentlicht: (2025)
3D Scene Generation: A Survey
von: Wen, Beichen, et al.
Veröffentlicht: (2025)
von: Wen, Beichen, et al.
Veröffentlicht: (2025)
Compositional Generative Model of Unbounded 4D Cities
von: Xie, Haozhe, et al.
Veröffentlicht: (2025)
von: Xie, Haozhe, et al.
Veröffentlicht: (2025)
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
von: Cao, Yukang, et al.
Veröffentlicht: (2026)
von: Cao, Yukang, et al.
Veröffentlicht: (2026)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
von: Lv, Zhengyao, et al.
Veröffentlicht: (2025)
von: Lv, Zhengyao, et al.
Veröffentlicht: (2025)
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
von: Chen, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoxi, et al.
Veröffentlicht: (2025)
GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
von: Wu, Yuchen, et al.
Veröffentlicht: (2026)
von: Wu, Yuchen, et al.
Veröffentlicht: (2026)
MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation
von: Zhang, Xujie, et al.
Veröffentlicht: (2024)
von: Zhang, Xujie, et al.
Veröffentlicht: (2024)
Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
Pandora3D: A Comprehensive Framework for High-Quality 3D Shape and Texture Generation
von: Yang, Jiayu, et al.
Veröffentlicht: (2025)
von: Yang, Jiayu, et al.
Veröffentlicht: (2025)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
DreamGaussian4D: Generative 4D Gaussian Splatting
von: Ren, Jiawei, et al.
Veröffentlicht: (2023)
von: Ren, Jiawei, et al.
Veröffentlicht: (2023)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
Multi-Modal Prompt Learning on Blind Image Quality Assessment
von: Pan, Wensheng, et al.
Veröffentlicht: (2024)
von: Pan, Wensheng, et al.
Veröffentlicht: (2024)
Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization
von: Wang, Xiaohan, et al.
Veröffentlicht: (2025)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2025)
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
TUNI: Real-time RGB-T Semantic Segmentation with Unified Multi-Modal Feature Extraction and Cross-Modal Feature Fusion
von: Guo, Xiaodong, et al.
Veröffentlicht: (2025)
von: Guo, Xiaodong, et al.
Veröffentlicht: (2025)
BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessment
von: Zhou, Kanglei, et al.
Veröffentlicht: (2026)
von: Zhou, Kanglei, et al.
Veröffentlicht: (2026)
ShapeGen: Towards High-Quality 3D Shape Synthesis
von: Li, Yangguang, et al.
Veröffentlicht: (2025)
von: Li, Yangguang, et al.
Veröffentlicht: (2025)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors
von: Huang, Zihao, et al.
Veröffentlicht: (2026)
von: Huang, Zihao, et al.
Veröffentlicht: (2026)
Towards Personalized Multi-Modal MRI Synthesis across Heterogeneous Datasets
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
Reconstructing 4D Spatial Intelligence: A Survey
von: Cao, Yukang, et al.
Veröffentlicht: (2025)
von: Cao, Yukang, et al.
Veröffentlicht: (2025)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
von: Qiu, Haonan, et al.
Veröffentlicht: (2024)
von: Qiu, Haonan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PhysX-3D: Physical-Grounded 3D Asset Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025) -
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
von: Cao, Ziang, et al.
Veröffentlicht: (2025) -
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
von: Cao, Ziang, et al.
Veröffentlicht: (2026) -
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2024) -
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
von: Liu, Zixian, et al.
Veröffentlicht: (2026)