Taming Feed-forward Reconstruction Models as Latent Encoders for 3D Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wizadwongsa, Suttisak, Zhou, Jinfan, Li, Edward, Park, Jeong Joon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Spatiotemporal Data Augmentation
by: Zhou, Jinfan, et al.
Published: (2025)
by: Zhou, Jinfan, et al.
Published: (2025)
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
by: Yao, Jingfeng, et al.
Published: (2025)
by: Yao, Jingfeng, et al.
Published: (2025)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025)
by: Li, Edward, et al.
Published: (2025)
Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters
by: Guo, Zhiyang, et al.
Published: (2025)
by: Guo, Zhiyang, et al.
Published: (2025)
Light3R-SfM: Towards Feed-forward Structure-from-Motion
by: Elflein, Sven, et al.
Published: (2025)
by: Elflein, Sven, et al.
Published: (2025)
Taming Latent Diffusion Model for Neural Radiance Field Inpainting
by: Lin, Chieh Hubert, et al.
Published: (2024)
by: Lin, Chieh Hubert, et al.
Published: (2024)
Taming Mode Collapse in Score Distillation for Text-to-3D Generation
by: Wang, Peihao, et al.
Published: (2023)
by: Wang, Peihao, et al.
Published: (2023)
Tuning Just Enough: Lightweight Backdoor Attacks on Multi-Encoder Diffusion Models
by: Chen, Ziyuan, et al.
Published: (2026)
by: Chen, Ziyuan, et al.
Published: (2026)
The Emergence of Reproducibility and Generalizability in Diffusion Models
by: Zhang, Huijie, et al.
Published: (2023)
by: Zhang, Huijie, et al.
Published: (2023)
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
by: Bratulić, Jelena, et al.
Published: (2025)
by: Bratulić, Jelena, et al.
Published: (2025)
LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models
by: Xie, Mingyang, et al.
Published: (2026)
by: Xie, Mingyang, et al.
Published: (2026)
Bayesian Diffusion Models for 3D Shape Reconstruction
by: Xu, Haiyang, et al.
Published: (2024)
by: Xu, Haiyang, et al.
Published: (2024)
Understanding Generalization in Diffusion Distillation via Probability Flow Distance
by: Zhang, Huijie, et al.
Published: (2025)
by: Zhang, Huijie, et al.
Published: (2025)
Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models
by: Henderson, Paul, et al.
Published: (2024)
by: Henderson, Paul, et al.
Published: (2024)
Spatiotemporal Satellite Image Downscaling with Transfer Encoders and Autoregressive Generative Models
by: Xiang, Yang, et al.
Published: (2025)
by: Xiang, Yang, et al.
Published: (2025)
Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning
by: Xie, Desai, et al.
Published: (2023)
by: Xie, Desai, et al.
Published: (2023)
Speed3R: Sparse Feed-forward 3D Reconstruction Models
by: Ren, Weining, et al.
Published: (2026)
by: Ren, Weining, et al.
Published: (2026)
Learning Multimodal Latent Generative Models with Energy-Based Prior
by: Yuan, Shiyu, et al.
Published: (2024)
by: Yuan, Shiyu, et al.
Published: (2024)
Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
by: Ren, Yifei, et al.
Published: (2025)
by: Ren, Yifei, et al.
Published: (2025)
Fast Autoregressive Models for Continuous Latent Generation
by: Hang, Tiankai, et al.
Published: (2025)
by: Hang, Tiankai, et al.
Published: (2025)
CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model
by: Wang, Zhengyi, et al.
Published: (2024)
by: Wang, Zhengyi, et al.
Published: (2024)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation
by: Lin, Siyou, et al.
Published: (2026)
by: Lin, Siyou, et al.
Published: (2026)
HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
by: Yu, Zhiyuan, et al.
Published: (2025)
by: Yu, Zhiyuan, et al.
Published: (2025)
Terra: Explorable Native 3D World Model with Point Latents
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
KeyPointDiffuser: Unsupervised 3D Keypoint Learning via Latent Diffusion Models
by: Newbury, Rhys, et al.
Published: (2025)
by: Newbury, Rhys, et al.
Published: (2025)
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
by: Keetha, Nikhil, et al.
Published: (2025)
by: Keetha, Nikhil, et al.
Published: (2025)
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
by: Jiang, Yitong, et al.
Published: (2026)
by: Jiang, Yitong, et al.
Published: (2026)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
Taming Generative Diffusion Prior for Universal Blind Image Restoration
by: Tu, Siwei, et al.
Published: (2024)
by: Tu, Siwei, et al.
Published: (2024)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction
by: Nguyen, Thuan Hoang, et al.
Published: (2026)
by: Nguyen, Thuan Hoang, et al.
Published: (2026)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
by: Li, Kevin, et al.
Published: (2025)
by: Li, Kevin, et al.
Published: (2025)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
by: Attaiki, Souhaib, et al.
Published: (2024)
by: Attaiki, Souhaib, et al.
Published: (2024)
DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
by: Sariyildiz, Mert Bulent, et al.
Published: (2025)
by: Sariyildiz, Mert Bulent, et al.
Published: (2025)
RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation
by: Anciukevičius, Titas, et al.
Published: (2022)
by: Anciukevičius, Titas, et al.
Published: (2022)
Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
by: Liu, Shengqi, et al.
Published: (2024)
by: Liu, Shengqi, et al.
Published: (2024)
Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
by: Ren, Weining, et al.
Published: (2025)
by: Ren, Weining, et al.
Published: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
by: Jin, Jiachun, et al.
Published: (2026)
by: Jin, Jiachun, et al.
Published: (2026)
Similar Items
-
Generative Spatiotemporal Data Augmentation
by: Zhou, Jinfan, et al.
Published: (2025) -
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
by: Yao, Jingfeng, et al.
Published: (2025) -
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025) -
Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters
by: Guo, Zhiyang, et al.
Published: (2025) -
Light3R-SfM: Towards Feed-forward Structure-from-Motion
by: Elflein, Sven, et al.
Published: (2025)