Template-Free Single-View 3D Human Digitalization with Diffusion-Guided LRM
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Weng, Zhenzhen, Liu, Jingyuan, Tan, Hao, Xu, Zhan, Zhou, Yang, Yeung-Levy, Serena, Yang, Jimei |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
par: Weng, Zhenzhen, et autres
Publié: (2023)
par: Weng, Zhenzhen, et autres
Publié: (2023)
Multi-Human Mesh Recovery with Transformers
par: Wang, Zeyu, et autres
Publié: (2024)
par: Wang, Zeyu, et autres
Publié: (2024)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
par: Wang, Zeyu, et autres
Publié: (2024)
par: Wang, Zeyu, et autres
Publié: (2024)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
par: Burgess, James, et autres
Publié: (2023)
par: Burgess, James, et autres
Publié: (2023)
LRM: Large Reconstruction Model for Single Image to 3D
par: Hong, Yicong, et autres
Publié: (2023)
par: Hong, Yicong, et autres
Publié: (2023)
Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera
par: Heo, Jaewoo, et autres
Publié: (2024)
par: Heo, Jaewoo, et autres
Publié: (2024)
RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets
par: Liu, Isabella, et autres
Publié: (2025)
par: Liu, Isabella, et autres
Publié: (2025)
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
par: Ma, Ziqiao, et autres
Publié: (2025)
par: Ma, Ziqiao, et autres
Publié: (2025)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
par: Heo, Jaewoo, et autres
Publié: (2024)
par: Heo, Jaewoo, et autres
Publié: (2024)
GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
par: Zhang, Kai, et autres
Publié: (2024)
par: Zhang, Kai, et autres
Publié: (2024)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
par: Bravo-Sánchez, Laura, et autres
Publié: (2024)
par: Bravo-Sánchez, Laura, et autres
Publié: (2024)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
par: Endo, Mark, et autres
Publié: (2025)
par: Endo, Mark, et autres
Publié: (2025)
iLRM: An Iterative Large 3D Reconstruction Model
par: Kang, Gyeongjin, et autres
Publié: (2025)
par: Kang, Gyeongjin, et autres
Publié: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
par: Gu, Jeffrey, et autres
Publié: (2025)
par: Gu, Jeffrey, et autres
Publié: (2025)
tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction
par: Wang, Chen, et autres
Publié: (2026)
par: Wang, Chen, et autres
Publié: (2026)
Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction
par: Ziwen, Chen, et autres
Publié: (2025)
par: Ziwen, Chen, et autres
Publié: (2025)
MeshLRM: Large Reconstruction Model for High-Quality Meshes
par: Wei, Xinyue, et autres
Publié: (2024)
par: Wei, Xinyue, et autres
Publié: (2024)
LRM-Zero: Training Large Reconstruction Models with Synthesized Data
par: Xie, Desai, et autres
Publié: (2024)
par: Xie, Desai, et autres
Publié: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
par: Aklilu, Josiah, et autres
Publié: (2024)
par: Aklilu, Josiah, et autres
Publié: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
par: Endo, Mark, et autres
Publié: (2024)
par: Endo, Mark, et autres
Publié: (2024)
ControLRM: Fast and Controllable 3D Generation via Large Reconstruction Model
par: Xu, Hongbin, et autres
Publié: (2024)
par: Xu, Hongbin, et autres
Publié: (2024)
FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
par: Jiang, Chenhan, et autres
Publié: (2026)
par: Jiang, Chenhan, et autres
Publié: (2026)
Long-LRM: Long-sequence Large Reconstruction Model for Wide-coverage Gaussian Splats
par: Ziwen, Chen, et autres
Publié: (2024)
par: Ziwen, Chen, et autres
Publié: (2024)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
par: Sui, Elaine, et autres
Publié: (2024)
par: Sui, Elaine, et autres
Publié: (2024)
ActAnywhere: Subject-Aware Video Background Generation
par: Pan, Boxiao, et autres
Publié: (2024)
par: Pan, Boxiao, et autres
Publié: (2024)
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
par: Jin, Yudong, et autres
Publié: (2025)
par: Jin, Yudong, et autres
Publié: (2025)
RelitLRM: Generative Relightable Radiance for Large Reconstruction Models
par: Zhang, Tianyuan, et autres
Publié: (2024)
par: Zhang, Tianyuan, et autres
Publié: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
par: Zhang, Yuhui, et autres
Publié: (2024)
par: Zhang, Yuhui, et autres
Publié: (2024)
HumanGif: Single-View Human Diffusion with Generative Prior
par: Hu, Shoukang, et autres
Publié: (2025)
par: Hu, Shoukang, et autres
Publié: (2025)
Anny-Fit: All-Age Human Mesh Recovery
par: Bravo-Sánchez, Laura, et autres
Publié: (2026)
par: Bravo-Sánchez, Laura, et autres
Publié: (2026)
ViewFormer: Exploring Spatiotemporal Modeling for Multi-View 3D Occupancy Perception via View-Guided Transformers
par: Li, Jinke, et autres
Publié: (2024)
par: Li, Jinke, et autres
Publié: (2024)
KaoLRM: Repurposing Pre-trained Large Reconstruction Models for Parametric 3D Face Reconstruction
par: Zhu, Qingtian, et autres
Publié: (2026)
par: Zhu, Qingtian, et autres
Publié: (2026)
GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction
par: Mu, Yuxuan, et autres
Publié: (2024)
par: Mu, Yuxuan, et autres
Publié: (2024)
Synergistic Global-space Camera and Human Reconstruction from Videos
par: Zhao, Yizhou, et autres
Publié: (2024)
par: Zhao, Yizhou, et autres
Publié: (2024)
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
par: Lin, Chieh Hubert, et autres
Publié: (2025)
par: Lin, Chieh Hubert, et autres
Publié: (2025)
Human Video Generation from a Single Image with 3D Pose and View Control
par: Wang, Tiantian, et autres
Publié: (2026)
par: Wang, Tiantian, et autres
Publié: (2026)
GeoFusionLRM: Geometry-Aware Self-Correction for Consistent 3D Reconstruction
par: Yildirim, Ahmet Burak, et autres
Publié: (2026)
par: Yildirim, Ahmet Burak, et autres
Publié: (2026)
Light Field Diffusion for Single-View Novel View Synthesis
par: Xiong, Yifeng, et autres
Publié: (2023)
par: Xiong, Yifeng, et autres
Publié: (2023)
EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head
par: He, Qianyun, et autres
Publié: (2024)
par: He, Qianyun, et autres
Publié: (2024)
X-LRM: X-ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One Second
par: Zhang, Guofeng, et autres
Publié: (2025)
par: Zhang, Guofeng, et autres
Publié: (2025)
Documents similaires
-
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
par: Weng, Zhenzhen, et autres
Publié: (2023) -
Multi-Human Mesh Recovery with Transformers
par: Wang, Zeyu, et autres
Publié: (2024) -
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
par: Wang, Zeyu, et autres
Publié: (2024) -
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
par: Burgess, James, et autres
Publié: (2023) -
LRM: Large Reconstruction Model for Single Image to 3D
par: Hong, Yicong, et autres
Publié: (2023)