MagicDrive: Street View Generation with Diverse 3D Geometry Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Ruiyuan, Chen, Kai, Xie, Enze, Hong, Lanqing, Li, Zhenguo, Yeung, Dit-Yan, Xu, Qiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024)
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024)
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024)
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
di: Chen, Kai, et al.
Pubblicazione: (2023)
di: Chen, Kai, et al.
Pubblicazione: (2023)
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
di: Wang, Yibo, et al.
Pubblicazione: (2024)
di: Wang, Yibo, et al.
Pubblicazione: (2024)
Mixed Autoencoder for Self-supervised Visual Representation Learning
di: Chen, Kai, et al.
Pubblicazione: (2023)
di: Chen, Kai, et al.
Pubblicazione: (2023)
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
di: Li, Pengxiang, et al.
Pubblicazione: (2023)
di: Li, Pengxiang, et al.
Pubblicazione: (2023)
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
di: Zhao, Yuyang, et al.
Pubblicazione: (2023)
di: Zhao, Yuyang, et al.
Pubblicazione: (2023)
Animate124: Animating One Image to 4D Dynamic Scene
di: Zhao, Yuyang, et al.
Pubblicazione: (2023)
di: Zhao, Yuyang, et al.
Pubblicazione: (2023)
ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
di: Chen, Kai, et al.
Pubblicazione: (2025)
di: Chen, Kai, et al.
Pubblicazione: (2025)
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
di: Chen, Kai, et al.
Pubblicazione: (2024)
di: Chen, Kai, et al.
Pubblicazione: (2024)
Implicit Concept Removal of Diffusion Models
di: Liu, Zhili, et al.
Pubblicazione: (2023)
di: Liu, Zhili, et al.
Pubblicazione: (2023)
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
di: Gou, Yunhao, et al.
Pubblicazione: (2024)
di: Gou, Yunhao, et al.
Pubblicazione: (2024)
CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
di: Jiang, Chenhan, et al.
Pubblicazione: (2025)
di: Jiang, Chenhan, et al.
Pubblicazione: (2025)
FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
di: Jiang, Chenhan, et al.
Pubblicazione: (2026)
di: Jiang, Chenhan, et al.
Pubblicazione: (2026)
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs
di: Zhong, Yingji, et al.
Pubblicazione: (2024)
di: Zhong, Yingji, et al.
Pubblicazione: (2024)
JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation
di: Jiang, Chenhan, et al.
Pubblicazione: (2024)
di: Jiang, Chenhan, et al.
Pubblicazione: (2024)
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
di: Wang, Tianqi, et al.
Pubblicazione: (2024)
di: Wang, Tianqi, et al.
Pubblicazione: (2024)
Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning
di: Gou, Yunhao, et al.
Pubblicazione: (2023)
di: Gou, Yunhao, et al.
Pubblicazione: (2023)
SplatMesh: Interactive 3D Segmentation and Editing Using Mesh-Based Gaussian Splatting
di: Zhou, Kaichen, et al.
Pubblicazione: (2023)
di: Zhou, Kaichen, et al.
Pubblicazione: (2023)
DreamDrive: Generative 4D Scene Modeling from Street View Images
di: Mao, Jiageng, et al.
Pubblicazione: (2024)
di: Mao, Jiageng, et al.
Pubblicazione: (2024)
Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views
di: Zhong, Yingji, et al.
Pubblicazione: (2025)
di: Zhong, Yingji, et al.
Pubblicazione: (2025)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
di: Cao, Yang, et al.
Pubblicazione: (2026)
di: Cao, Yang, et al.
Pubblicazione: (2026)
TransformMix: Learning Transformation and Mixing Strategies from Data
di: Cheung, Tsz-Him, et al.
Pubblicazione: (2024)
di: Cheung, Tsz-Him, et al.
Pubblicazione: (2024)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
di: Yan, Yunzhi, et al.
Pubblicazione: (2024)
di: Yan, Yunzhi, et al.
Pubblicazione: (2024)
Text2Street: Controllable Text-to-image Generation for Street Views
di: Su, Jinming, et al.
Pubblicazione: (2024)
di: Su, Jinming, et al.
Pubblicazione: (2024)
SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs
di: Li, Leheng, et al.
Pubblicazione: (2024)
di: Li, Leheng, et al.
Pubblicazione: (2024)
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
di: Wu, Junjie, et al.
Pubblicazione: (2024)
di: Wu, Junjie, et al.
Pubblicazione: (2024)
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
di: Chen, Junsong, et al.
Pubblicazione: (2024)
di: Chen, Junsong, et al.
Pubblicazione: (2024)
Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection
di: Chen, Zhili, et al.
Pubblicazione: (2024)
di: Chen, Zhili, et al.
Pubblicazione: (2024)
MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement
di: He, Xu, et al.
Pubblicazione: (2024)
di: He, Xu, et al.
Pubblicazione: (2024)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
di: Wu, Zehuan, et al.
Pubblicazione: (2024)
di: Wu, Zehuan, et al.
Pubblicazione: (2024)
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
di: Liu, Zhili, et al.
Pubblicazione: (2024)
di: Liu, Zhili, et al.
Pubblicazione: (2024)
Learning 3D Persistent Embodied World Models
di: Zhou, Siyuan, et al.
Pubblicazione: (2025)
di: Zhou, Siyuan, et al.
Pubblicazione: (2025)
Street-View Image Generation from a Bird's-Eye View Layout
di: Swerdlow, Alexander, et al.
Pubblicazione: (2023)
di: Swerdlow, Alexander, et al.
Pubblicazione: (2023)
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
di: Huang, Kaiyi, et al.
Pubblicazione: (2023)
di: Huang, Kaiyi, et al.
Pubblicazione: (2023)
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
di: Wang, Zhenyu, et al.
Pubblicazione: (2024)
di: Wang, Zhenyu, et al.
Pubblicazione: (2024)
Magic-Boost: Boost 3D Generation with Multi-View Conditioned Diffusion
di: Yang, Fan, et al.
Pubblicazione: (2024)
di: Yang, Fan, et al.
Pubblicazione: (2024)
DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation
di: Yan, Junkai, et al.
Pubblicazione: (2024)
di: Yan, Junkai, et al.
Pubblicazione: (2024)
SVIA: A Street View Image Anonymization Framework for Self-Driving Applications
di: Liu, Dongyu, et al.
Pubblicazione: (2025)
di: Liu, Dongyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024) -
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024) -
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
di: Chen, Kai, et al.
Pubblicazione: (2023) -
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
di: Wang, Yibo, et al.
Pubblicazione: (2024) -
Mixed Autoencoder for Self-supervised Visual Representation Learning
di: Chen, Kai, et al.
Pubblicazione: (2023)