Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Go, Hyojun, Narnhofer, Dominik, Bhat, Goutam, Truong, Prune, Tombari, Federico, Schindler, Konrad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stitched Value Model for Diffusion Alignment
von: Go, Hyojun, et al.
Veröffentlicht: (2026)
von: Go, Hyojun, et al.
Veröffentlicht: (2026)
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
von: Metzger, Nando, et al.
Veröffentlicht: (2025)
von: Metzger, Nando, et al.
Veröffentlicht: (2025)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
Understanding, Accelerating, and Improving MeanFlow Training
von: Kim, Jin-Young, et al.
Veröffentlicht: (2025)
von: Kim, Jin-Young, et al.
Veröffentlicht: (2025)
Continuous Space-Time Video Super-Resolution with 3D Fourier Fields
von: Becker, Alexander, et al.
Veröffentlicht: (2025)
von: Becker, Alexander, et al.
Veröffentlicht: (2025)
Generating Human Motion Videos using a Cascaded Text-to-Video Framework
von: Nam, Hyelin, et al.
Veröffentlicht: (2025)
von: Nam, Hyelin, et al.
Veröffentlicht: (2025)
One2Any: One-Reference 6D Pose Estimation for Any Object
von: Liu, Mengya, et al.
Veröffentlicht: (2025)
von: Liu, Mengya, et al.
Veröffentlicht: (2025)
FlowSDF: Flow Matching for Medical Image Segmentation Using Distance Transforms
von: Bogensperger, Lea, et al.
Veröffentlicht: (2024)
von: Bogensperger, Lea, et al.
Veröffentlicht: (2024)
Video Depth without Video Models
von: Ke, Bingxin, et al.
Veröffentlicht: (2024)
von: Ke, Bingxin, et al.
Veröffentlicht: (2024)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
Unified Panoramic Geometry Estimation via Multi-View Foundation Models
von: Bozic, Vukasin, et al.
Veröffentlicht: (2026)
von: Bozic, Vukasin, et al.
Veröffentlicht: (2026)
CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
von: Kong, Xin, et al.
Veröffentlicht: (2025)
von: Kong, Xin, et al.
Veröffentlicht: (2025)
Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
von: Di Lorenzo, Gaia, et al.
Veröffentlicht: (2025)
von: Di Lorenzo, Gaia, et al.
Veröffentlicht: (2025)
Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D Environments
von: Zhu, Liyuan, et al.
Veröffentlicht: (2023)
von: Zhu, Liyuan, et al.
Veröffentlicht: (2023)
CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation
von: Kalischek, Nikolai, et al.
Veröffentlicht: (2025)
von: Kalischek, Nikolai, et al.
Veröffentlicht: (2025)
HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape Generation
von: Leng, Zhiying, et al.
Veröffentlicht: (2024)
von: Leng, Zhiying, et al.
Veröffentlicht: (2024)
Solving Inverse Problems with FLAIR
von: Erbach, Julius, et al.
Veröffentlicht: (2025)
von: Erbach, Julius, et al.
Veröffentlicht: (2025)
AnyUp: Universal Feature Upsampling
von: Wimmer, Thomas, et al.
Veröffentlicht: (2025)
von: Wimmer, Thomas, et al.
Veröffentlicht: (2025)
LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
MOHO: Learning Single-view Hand-held Object Reconstruction with Multi-view Occlusion-Aware Supervision
von: Zhang, Chenyangguang, et al.
Veröffentlicht: (2023)
von: Zhang, Chenyangguang, et al.
Veröffentlicht: (2023)
Gaussians-to-Life: Text-Driven Animation of 3D Gaussian Splatting Scenes
von: Wimmer, Thomas, et al.
Veröffentlicht: (2024)
von: Wimmer, Thomas, et al.
Veröffentlicht: (2024)
Thera: Aliasing-Free Arbitrary-Scale Super-Resolution with Neural Heat Fields
von: Becker, Alexander, et al.
Veröffentlicht: (2023)
von: Becker, Alexander, et al.
Veröffentlicht: (2023)
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis
von: Go, Hyojun, et al.
Veröffentlicht: (2024)
von: Go, Hyojun, et al.
Veröffentlicht: (2024)
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
von: Sun, Fan-Yun, et al.
Veröffentlicht: (2024)
von: Sun, Fan-Yun, et al.
Veröffentlicht: (2024)
Text-Conditioned Resampler For Long Form Video Understanding
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
Epipolar Geometry Improves Video Generation Models
von: Kupyn, Orest, et al.
Veröffentlicht: (2025)
von: Kupyn, Orest, et al.
Veröffentlicht: (2025)
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
von: Park, Byeongjun, et al.
Veröffentlicht: (2022)
von: Park, Byeongjun, et al.
Veröffentlicht: (2022)
MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly Detection
von: Micorek, Jakub, et al.
Veröffentlicht: (2024)
von: Micorek, Jakub, et al.
Veröffentlicht: (2024)
MUSt3R: Multi-view Network for Stereo 3D Reconstruction
von: Cabon, Yohann, et al.
Veröffentlicht: (2025)
von: Cabon, Yohann, et al.
Veröffentlicht: (2025)
SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering
von: Park, Byeongjun, et al.
Veröffentlicht: (2025)
von: Park, Byeongjun, et al.
Veröffentlicht: (2025)
SING3R-SLAM: Submap-based Indoor Monocular Gaussian SLAM with 3D Reconstruction Priors
von: Li, Kunyi, et al.
Veröffentlicht: (2025)
von: Li, Kunyi, et al.
Veröffentlicht: (2025)
UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections
von: Wang, Fangjinhua, et al.
Veröffentlicht: (2023)
von: Wang, Fangjinhua, et al.
Veröffentlicht: (2023)
Video Perception Models for 3D Scene Synthesis
von: Huang, Rui, et al.
Veröffentlicht: (2025)
von: Huang, Rui, et al.
Veröffentlicht: (2025)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
von: Kwon, Soonwoo, et al.
Veröffentlicht: (2025)
von: Kwon, Soonwoo, et al.
Veröffentlicht: (2025)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
von: Shahbazi, Mohamad, et al.
Veröffentlicht: (2024)
von: Shahbazi, Mohamad, et al.
Veröffentlicht: (2024)
RaNeuS: Ray-adaptive Neural Surface Reconstruction
von: Wang, Yida, et al.
Veröffentlicht: (2024)
von: Wang, Yida, et al.
Veröffentlicht: (2024)
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
von: Cai, Xiao, et al.
Veröffentlicht: (2024)
von: Cai, Xiao, et al.
Veröffentlicht: (2024)
Video Self-Stitching Graph Network for Temporal Action Localization
von: Zhao, Chen, et al.
Veröffentlicht: (2020)
von: Zhao, Chen, et al.
Veröffentlicht: (2020)
Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation
von: Kogure, Hina, et al.
Veröffentlicht: (2026)
von: Kogure, Hina, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Stitched Value Model for Diffusion Alignment
von: Go, Hyojun, et al.
Veröffentlicht: (2026) -
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
von: Metzger, Nando, et al.
Veröffentlicht: (2025) -
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025) -
Understanding, Accelerating, and Improving MeanFlow Training
von: Kim, Jin-Young, et al.
Veröffentlicht: (2025) -
Continuous Space-Time Video Super-Resolution with 3D Fourier Fields
von: Becker, Alexander, et al.
Veröffentlicht: (2025)