Taming Feed-forward Reconstruction Models as Latent Encoders for 3D Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wizadwongsa, Suttisak, Zhou, Jinfan, Li, Edward, Park, Jeong Joon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative Spatiotemporal Data Augmentation
von: Zhou, Jinfan, et al.
Veröffentlicht: (2025)
von: Zhou, Jinfan, et al.
Veröffentlicht: (2025)
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
von: Yao, Jingfeng, et al.
Veröffentlicht: (2025)
von: Yao, Jingfeng, et al.
Veröffentlicht: (2025)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
von: Li, Edward, et al.
Veröffentlicht: (2025)
von: Li, Edward, et al.
Veröffentlicht: (2025)
Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters
von: Guo, Zhiyang, et al.
Veröffentlicht: (2025)
von: Guo, Zhiyang, et al.
Veröffentlicht: (2025)
Light3R-SfM: Towards Feed-forward Structure-from-Motion
von: Elflein, Sven, et al.
Veröffentlicht: (2025)
von: Elflein, Sven, et al.
Veröffentlicht: (2025)
Taming Latent Diffusion Model for Neural Radiance Field Inpainting
von: Lin, Chieh Hubert, et al.
Veröffentlicht: (2024)
von: Lin, Chieh Hubert, et al.
Veröffentlicht: (2024)
Taming Mode Collapse in Score Distillation for Text-to-3D Generation
von: Wang, Peihao, et al.
Veröffentlicht: (2023)
von: Wang, Peihao, et al.
Veröffentlicht: (2023)
Tuning Just Enough: Lightweight Backdoor Attacks on Multi-Encoder Diffusion Models
von: Chen, Ziyuan, et al.
Veröffentlicht: (2026)
von: Chen, Ziyuan, et al.
Veröffentlicht: (2026)
The Emergence of Reproducibility and Generalizability in Diffusion Models
von: Zhang, Huijie, et al.
Veröffentlicht: (2023)
von: Zhang, Huijie, et al.
Veröffentlicht: (2023)
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models
von: Xie, Mingyang, et al.
Veröffentlicht: (2026)
von: Xie, Mingyang, et al.
Veröffentlicht: (2026)
Bayesian Diffusion Models for 3D Shape Reconstruction
von: Xu, Haiyang, et al.
Veröffentlicht: (2024)
von: Xu, Haiyang, et al.
Veröffentlicht: (2024)
Understanding Generalization in Diffusion Distillation via Probability Flow Distance
von: Zhang, Huijie, et al.
Veröffentlicht: (2025)
von: Zhang, Huijie, et al.
Veröffentlicht: (2025)
Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models
von: Henderson, Paul, et al.
Veröffentlicht: (2024)
von: Henderson, Paul, et al.
Veröffentlicht: (2024)
Spatiotemporal Satellite Image Downscaling with Transfer Encoders and Autoregressive Generative Models
von: Xiang, Yang, et al.
Veröffentlicht: (2025)
von: Xiang, Yang, et al.
Veröffentlicht: (2025)
Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning
von: Xie, Desai, et al.
Veröffentlicht: (2023)
von: Xie, Desai, et al.
Veröffentlicht: (2023)
Speed3R: Sparse Feed-forward 3D Reconstruction Models
von: Ren, Weining, et al.
Veröffentlicht: (2026)
von: Ren, Weining, et al.
Veröffentlicht: (2026)
Learning Multimodal Latent Generative Models with Energy-Based Prior
von: Yuan, Shiyu, et al.
Veröffentlicht: (2024)
von: Yuan, Shiyu, et al.
Veröffentlicht: (2024)
Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
von: Ren, Yifei, et al.
Veröffentlicht: (2025)
von: Ren, Yifei, et al.
Veröffentlicht: (2025)
Fast Autoregressive Models for Continuous Latent Generation
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model
von: Wang, Zhengyi, et al.
Veröffentlicht: (2024)
von: Wang, Zhengyi, et al.
Veröffentlicht: (2024)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation
von: Lin, Siyou, et al.
Veröffentlicht: (2026)
von: Lin, Siyou, et al.
Veröffentlicht: (2026)
HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2025)
Terra: Explorable Native 3D World Model with Point Latents
von: Huang, Yuanhui, et al.
Veröffentlicht: (2025)
von: Huang, Yuanhui, et al.
Veröffentlicht: (2025)
KeyPointDiffuser: Unsupervised 3D Keypoint Learning via Latent Diffusion Models
von: Newbury, Rhys, et al.
Veröffentlicht: (2025)
von: Newbury, Rhys, et al.
Veröffentlicht: (2025)
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
von: Keetha, Nikhil, et al.
Veröffentlicht: (2025)
von: Keetha, Nikhil, et al.
Veröffentlicht: (2025)
Unifying Specialized Visual Encoders for Video Language Models
von: Chung, Jihoon, et al.
Veröffentlicht: (2025)
von: Chung, Jihoon, et al.
Veröffentlicht: (2025)
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
von: Jiang, Yitong, et al.
Veröffentlicht: (2026)
von: Jiang, Yitong, et al.
Veröffentlicht: (2026)
Taming Generative Diffusion Prior for Universal Blind Image Restoration
von: Tu, Siwei, et al.
Veröffentlicht: (2024)
von: Tu, Siwei, et al.
Veröffentlicht: (2024)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction
von: Nguyen, Thuan Hoang, et al.
Veröffentlicht: (2026)
von: Nguyen, Thuan Hoang, et al.
Veröffentlicht: (2026)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
von: Li, Kevin, et al.
Veröffentlicht: (2025)
von: Li, Kevin, et al.
Veröffentlicht: (2025)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
von: Attaiki, Souhaib, et al.
Veröffentlicht: (2024)
von: Attaiki, Souhaib, et al.
Veröffentlicht: (2024)
DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
von: Sariyildiz, Mert Bulent, et al.
Veröffentlicht: (2025)
von: Sariyildiz, Mert Bulent, et al.
Veröffentlicht: (2025)
RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation
von: Anciukevičius, Titas, et al.
Veröffentlicht: (2022)
von: Anciukevičius, Titas, et al.
Veröffentlicht: (2022)
Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
von: Liu, Shengqi, et al.
Veröffentlicht: (2024)
von: Liu, Shengqi, et al.
Veröffentlicht: (2024)
Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
von: Ren, Weining, et al.
Veröffentlicht: (2025)
von: Ren, Weining, et al.
Veröffentlicht: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
von: Nilaksh, et al.
Veröffentlicht: (2026)
von: Nilaksh, et al.
Veröffentlicht: (2026)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Generative Spatiotemporal Data Augmentation
von: Zhou, Jinfan, et al.
Veröffentlicht: (2025) -
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
von: Yao, Jingfeng, et al.
Veröffentlicht: (2025) -
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
von: Li, Edward, et al.
Veröffentlicht: (2025) -
Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters
von: Guo, Zhiyang, et al.
Veröffentlicht: (2025) -
Light3R-SfM: Towards Feed-forward Structure-from-Motion
von: Elflein, Sven, et al.
Veröffentlicht: (2025)