MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Chenjie, Yu, Chaohui, Liu, Shang, Wang, Fan, Xue, Xiangyang, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image
by: Ren, Xinlin, et al.
Published: (2024)
by: Ren, Xinlin, et al.
Published: (2024)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
by: Cao, Chenjie, et al.
Published: (2025)
by: Cao, Chenjie, et al.
Published: (2025)
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion
by: Jiang, Yanqin, et al.
Published: (2024)
by: Jiang, Yanqin, et al.
Published: (2024)
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
by: Liu, Shang, et al.
Published: (2025)
by: Liu, Shang, et al.
Published: (2025)
Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
by: Wang, Yikai, et al.
Published: (2023)
by: Wang, Yikai, et al.
Published: (2023)
AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks
by: Xie, Ming, et al.
Published: (2025)
by: Xie, Ming, et al.
Published: (2025)
MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
VCD-Texture: Variance Alignment based 3D-2D Co-Denoising for Text-Guided Texturing
by: Liu, Shang, et al.
Published: (2024)
by: Liu, Shang, et al.
Published: (2024)
Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
by: Wang, Yikai, et al.
Published: (2026)
by: Wang, Yikai, et al.
Published: (2026)
MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models
by: Fang, Shaoheng, et al.
Published: (2025)
by: Fang, Shaoheng, et al.
Published: (2025)
Repositioning the Subject within Image
by: Wang, Yikai, et al.
Published: (2024)
by: Wang, Yikai, et al.
Published: (2024)
LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion
by: Zhang, Yisu, et al.
Published: (2025)
by: Zhang, Yisu, et al.
Published: (2025)
LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model
by: Cao, Chenjie, et al.
Published: (2023)
by: Cao, Chenjie, et al.
Published: (2023)
SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer
by: Wu, Zijie, et al.
Published: (2024)
by: Wu, Zijie, et al.
Published: (2024)
SparseGrasp: Robotic Grasping via 3D Semantic Gaussian Splatting from Sparse Multi-View RGB Images
by: Yu, Junqiu, et al.
Published: (2024)
by: Yu, Junqiu, et al.
Published: (2024)
Sub-Image Recapture for Multi-View 3D Reconstruction
by: Wang, Yanwei
Published: (2025)
by: Wang, Yanwei
Published: (2025)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
by: Wei, Min, et al.
Published: (2025)
by: Wei, Min, et al.
Published: (2025)
Enhancing Video Inpainting with Aligned Frame Interval Guidance
by: Xie, Ming, et al.
Published: (2025)
by: Xie, Ming, et al.
Published: (2025)
Beyond 'Templates': Category-Agnostic Object Pose, Size, and Shape Estimation from a Single View
by: Zhang, Jinyu, et al.
Published: (2025)
by: Zhang, Jinyu, et al.
Published: (2025)
Any-to-3D Generation via Hybrid Diffusion Supervision
by: Fan, Yijun, et al.
Published: (2024)
by: Fan, Yijun, et al.
Published: (2024)
AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation
by: Wu, Zijie, et al.
Published: (2026)
by: Wu, Zijie, et al.
Published: (2026)
AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
by: Wu, Zijie, et al.
Published: (2025)
by: Wu, Zijie, et al.
Published: (2025)
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
by: Chen, Yutian, et al.
Published: (2026)
by: Chen, Yutian, et al.
Published: (2026)
Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D Prior
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States
by: Wang, Qizao, et al.
Published: (2024)
by: Wang, Qizao, et al.
Published: (2024)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
by: Liu, Zhenyang, et al.
Published: (2026)
by: Liu, Zhenyang, et al.
Published: (2026)
CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image
by: Huang, Jingshun, et al.
Published: (2025)
by: Huang, Jingshun, et al.
Published: (2025)
SCOOP'D: Learning Mixed-Liquid-Solid Scooping via Sim2Real Generative Policy
by: Wang, Kuanning, et al.
Published: (2025)
by: Wang, Kuanning, et al.
Published: (2025)
Dereflection Any Image with Diffusion Priors and Diversified Data
by: Hu, Jichen, et al.
Published: (2025)
by: Hu, Jichen, et al.
Published: (2025)
RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base
by: Wang, Kuanning, et al.
Published: (2025)
by: Wang, Kuanning, et al.
Published: (2025)
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
by: Chen, Yabo, et al.
Published: (2024)
by: Chen, Yabo, et al.
Published: (2024)
Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
by: Li, Shanshan, et al.
Published: (2025)
by: Li, Shanshan, et al.
Published: (2025)
Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
by: Yang, Feng, et al.
Published: (2025)
by: Yang, Feng, et al.
Published: (2025)
SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
by: Zhang, Lingyun, et al.
Published: (2025)
by: Zhang, Lingyun, et al.
Published: (2025)
RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation
by: Huang, Hanzhuo, et al.
Published: (2026)
by: Huang, Hanzhuo, et al.
Published: (2026)
3D Skew-Normal Splatting
by: Wu, Xiangru, et al.
Published: (2026)
by: Wu, Xiangru, et al.
Published: (2026)
V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
by: Pan, Jiancheng, et al.
Published: (2025)
by: Pan, Jiancheng, et al.
Published: (2025)
Similar Items
-
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
by: Cao, Chenjie, et al.
Published: (2024) -
Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image
by: Ren, Xinlin, et al.
Published: (2024) -
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
by: Cao, Chenjie, et al.
Published: (2025) -
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion
by: Jiang, Yanqin, et al.
Published: (2024) -
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
by: Liu, Shang, et al.
Published: (2025)