IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuan, Bai, Ziqian, Tan, Feitong, Cui, Zhaopeng, Fanello, Sean, Zhang, Yinda |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table Blendshapes
by: Bai, Ziqian, et al.
Published: (2024)
by: Bai, Ziqian, et al.
Published: (2024)
SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix
by: Dai, Peng, et al.
Published: (2024)
by: Dai, Peng, et al.
Published: (2024)
One2Avatar: Generative Implicit Head Avatar For Few-shot User Adaptation
by: Yu, Zhixuan, et al.
Published: (2024)
by: Yu, Zhixuan, et al.
Published: (2024)
S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
by: Dai, Peng, et al.
Published: (2025)
by: Dai, Peng, et al.
Published: (2025)
Talking Together: Synthesizing Co-Located 3D Conversations from Audio
by: Shan, Mengyi, et al.
Published: (2026)
by: Shan, Mengyi, et al.
Published: (2026)
LightAvatar: Efficient Head Avatar as Dynamic Neural Light Field
by: Wang, Huan, et al.
Published: (2024)
by: Wang, Huan, et al.
Published: (2024)
Archon: A Unified Multimodal Model for Holistic Digital Human Generation
by: Bao, Chong, et al.
Published: (2026)
by: Bao, Chong, et al.
Published: (2026)
SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model
by: Tan, Weipeng, et al.
Published: (2024)
by: Tan, Weipeng, et al.
Published: (2024)
EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
by: Fei, Zhengcong, et al.
Published: (2025)
by: Fei, Zhengcong, et al.
Published: (2025)
Real-time 3D-aware Portrait Video Relighting
by: Cai, Ziqi, et al.
Published: (2024)
by: Cai, Ziqi, et al.
Published: (2024)
InsTaG: Learning Personalized 3D Talking Head from Few-Second Video
by: Li, Jiahe, et al.
Published: (2025)
by: Li, Jiahe, et al.
Published: (2025)
High-Fidelity Relightable Monocular Portrait Animation with Lighting-Controllable Video Diffusion Model
by: Guo, Mingtao, et al.
Published: (2025)
by: Guo, Mingtao, et al.
Published: (2025)
Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2025)
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2025)
Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
by: Bai, Haoran, et al.
Published: (2025)
by: Bai, Haoran, et al.
Published: (2025)
Rig3DGS: Creating Controllable Portraits from Casual Monocular Videos
by: Rivero, Alfredo, et al.
Published: (2024)
by: Rivero, Alfredo, et al.
Published: (2024)
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
Latency-aware Road Anomaly Segmentation in Videos: A Photorealistic Dataset and New Metrics
by: Tian, Beiwen, et al.
Published: (2024)
by: Tian, Beiwen, et al.
Published: (2024)
Splat-Portrait: Generalizing Talking Heads with Gaussian Splatting
by: Shi, Tong, et al.
Published: (2026)
by: Shi, Tong, et al.
Published: (2026)
I'M HOI: Inertia-aware Monocular Capture of 3D Human-Object Interactions
by: Zhao, Chengfeng, et al.
Published: (2023)
by: Zhao, Chengfeng, et al.
Published: (2023)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
Monocular and Generalizable Gaussian Talking Head Animation
by: Gong, Shengjie, et al.
Published: (2025)
by: Gong, Shengjie, et al.
Published: (2025)
CFSynthesis: Controllable and Free-view 3D Human Video Synthesis
by: Cui, Liyuan, et al.
Published: (2024)
by: Cui, Liyuan, et al.
Published: (2024)
MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
by: Ye, Zhenhui, et al.
Published: (2024)
by: Ye, Zhenhui, et al.
Published: (2024)
GMTalker: Gaussian Mixture-based Audio-Driven Emotional Talking Video Portraits
by: Xia, Yibo, et al.
Published: (2023)
by: Xia, Yibo, et al.
Published: (2023)
Context-aware Talking Face Video Generation
by: Xuanyuan, Meidai, et al.
Published: (2024)
by: Xuanyuan, Meidai, et al.
Published: (2024)
RMAvatar: Photorealistic Human Avatar Reconstruction from Monocular Video Based on Rectified Mesh-embedded Gaussians
by: Peng, Sen, et al.
Published: (2025)
by: Peng, Sen, et al.
Published: (2025)
THEval. Evaluation Framework for Talking Head Video Generation
by: Quignon, Nabyl, et al.
Published: (2025)
by: Quignon, Nabyl, et al.
Published: (2025)
GeneAvatar: Generic Expression-Aware Volumetric Head Avatar Editing from a Single Image
by: Bao, Chong, et al.
Published: (2024)
by: Bao, Chong, et al.
Published: (2024)
PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing
by: Liang, Jiadong, et al.
Published: (2026)
by: Liang, Jiadong, et al.
Published: (2026)
DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos
by: Wu, Xiuzhe, et al.
Published: (2024)
by: Wu, Xiuzhe, et al.
Published: (2024)
GO-NeRF: Generating Objects in Neural Radiance Fields for Virtual Reality Content Creation
by: Dai, Peng, et al.
Published: (2024)
by: Dai, Peng, et al.
Published: (2024)
Anything in Any Scene: Photorealistic Video Object Insertion
by: Bai, Chen, et al.
Published: (2024)
by: Bai, Chen, et al.
Published: (2024)
4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models
by: Yu, Heng, et al.
Published: (2024)
by: Yu, Heng, et al.
Published: (2024)
Leveraging Avatar Fingerprinting: A Multi-Generator Photorealistic Talking-Head Public Database and Benchmark
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2026)
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2026)
SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
by: Qiu, Di, et al.
Published: (2025)
by: Qiu, Di, et al.
Published: (2025)
Learning Online Scale Transformation for Talking Head Video Generation
by: Hong, Fa-Ting, et al.
Published: (2024)
by: Hong, Fa-Ting, et al.
Published: (2024)
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
by: Wang, Haotian, et al.
Published: (2024)
by: Wang, Haotian, et al.
Published: (2024)
FAGhead: Fully Animate Gaussian Head from Monocular Videos
by: Xuan, Yixin, et al.
Published: (2024)
by: Xuan, Yixin, et al.
Published: (2024)
Similar Items
-
Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table Blendshapes
by: Bai, Ziqian, et al.
Published: (2024) -
SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix
by: Dai, Peng, et al.
Published: (2024) -
One2Avatar: Generative Implicit Head Avatar For Few-shot User Adaptation
by: Yu, Zhixuan, et al.
Published: (2024) -
S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
by: Dai, Peng, et al.
Published: (2025) -
Talking Together: Synthesizing Co-Located 3D Conversations from Audio
by: Shan, Mengyi, et al.
Published: (2026)