Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Chaolong, Yao, Kai, Yan, Yuyao, Jiang, Chenru, Zhao, Weiguang, Sun, Jie, Cheng, Guangliang, Zhang, Yifei, Dong, Bin, Huang, Kaizhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-Pose 3D Zero-Shot Learning: Benchmark and Challenges
by: Zhao, Weiguang, et al.
Published: (2023)
by: Zhao, Weiguang, et al.
Published: (2023)
3D-CDRGP: Towards Cross-Device Robotic Grasping Policy in 3D Open World
by: Zhao, Weiguang, et al.
Published: (2024)
by: Zhao, Weiguang, et al.
Published: (2024)
Consistency Diffusion Models for Single-Image 3D Reconstruction with Priors
by: Jiang, Chenru, et al.
Published: (2025)
by: Jiang, Chenru, et al.
Published: (2025)
From 2D Images to 3D Model:Weakly Supervised Multi-View Face Reconstruction with Deep Fusion
by: Zhao, Weiguang, et al.
Published: (2022)
by: Zhao, Weiguang, et al.
Published: (2022)
PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly Detection
by: Ye, Jianan, et al.
Published: (2024)
by: Ye, Jianan, et al.
Published: (2024)
BFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysis
by: Zhao, Weiguang, et al.
Published: (2025)
by: Zhao, Weiguang, et al.
Published: (2025)
Controllable Talking Face Generation by Implicit Facial Keypoints Editing
by: Zhao, Dong, et al.
Published: (2024)
by: Zhao, Dong, et al.
Published: (2024)
IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
by: Chen, Bo, et al.
Published: (2025)
by: Chen, Bo, et al.
Published: (2025)
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
by: Fei, Zhengcong, et al.
Published: (2025)
by: Fei, Zhengcong, et al.
Published: (2025)
FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait
by: Ki, Taekyung, et al.
Published: (2024)
by: Ki, Taekyung, et al.
Published: (2024)
FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis
by: Ni, Ziqi, et al.
Published: (2025)
by: Ni, Ziqi, et al.
Published: (2025)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024)
by: Nazarieh, Fatemeh, et al.
Published: (2024)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
LinguaLinker: Audio-Driven Portraits Animation with Implicit Facial Control Enhancement
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
X-Pose: Detecting Any Keypoints
by: Yang, Jie, et al.
Published: (2023)
by: Yang, Jie, et al.
Published: (2023)
READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
by: Wang, Haotian, et al.
Published: (2025)
by: Wang, Haotian, et al.
Published: (2025)
Towards Training-Free Open-World Classification with 3D Generative Models
by: Xia, Xinzhe, et al.
Published: (2025)
by: Xia, Xinzhe, et al.
Published: (2025)
RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
by: Du, Fangyu, et al.
Published: (2025)
by: Du, Fangyu, et al.
Published: (2025)
GMTalker: Gaussian Mixture-based Audio-Driven Emotional Talking Video Portraits
by: Xia, Yibo, et al.
Published: (2023)
by: Xia, Yibo, et al.
Published: (2023)
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
by: Salehi, Pegah, et al.
Published: (2024)
by: Salehi, Pegah, et al.
Published: (2024)
Portraits and Poses
Published: (2022)
Published: (2022)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
Diffusion Model Regularized Implicit Neural Representation for CT Metal Artifact Reduction
by: Wen, Jie, et al.
Published: (2025)
by: Wen, Jie, et al.
Published: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
by: Ling, Jun, et al.
Published: (2024)
by: Ling, Jun, et al.
Published: (2024)
Pareto-Guided Optimization for Uncertainty-Aware Medical Image Segmentation
by: Zhang, Jinming, et al.
Published: (2026)
by: Zhang, Jinming, et al.
Published: (2026)
RoboKeyGen: Robot Pose and Joint Angles Estimation via Diffusion-based 3D Keypoint Generation
by: Tian, Yang, et al.
Published: (2024)
by: Tian, Yang, et al.
Published: (2024)
Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation
by: Shi, Yifei, et al.
Published: (2026)
by: Shi, Yifei, et al.
Published: (2026)
SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model
by: Tan, Weipeng, et al.
Published: (2024)
by: Tan, Weipeng, et al.
Published: (2024)
FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
by: Wang, MengChao, et al.
Published: (2025)
by: Wang, MengChao, et al.
Published: (2025)
FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
by: Tang, Jiapeng, et al.
Published: (2025)
by: Tang, Jiapeng, et al.
Published: (2025)
TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
by: Javanmardi, Alireza, et al.
Published: (2025)
by: Javanmardi, Alireza, et al.
Published: (2025)
Spatially and Temporally Optimized Audio‐Driven Talking Face Generation
by: Biao Dong, et al.
Published: (2024)
by: Biao Dong, et al.
Published: (2024)
Towards Diverse and Efficient Audio Captioning via Diffusion Models
by: Xu, Manjie, et al.
Published: (2024)
by: Xu, Manjie, et al.
Published: (2024)
MK-Pose: Category-Level Object Pose Estimation via Multimodal-Based Keypoint Learning
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
by: Choi, Jeongsoo, et al.
Published: (2023)
by: Choi, Jeongsoo, et al.
Published: (2023)
Tele-Catch: Adaptive Teleoperation for Dexterous Dynamic 3D Object Catching
by: Zhao, Weiguang, et al.
Published: (2026)
by: Zhao, Weiguang, et al.
Published: (2026)
HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
by: Xu, Zunnan, et al.
Published: (2025)
by: Xu, Zunnan, et al.
Published: (2025)
DvD: Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion Model
by: Zhang, Weiguang, et al.
Published: (2025)
by: Zhang, Weiguang, et al.
Published: (2025)
FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
by: Wang, Mengchao, et al.
Published: (2025)
by: Wang, Mengchao, et al.
Published: (2025)
HOMER: Homography-Based Efficient Multi-view 3D Object Removal
by: Ni, Jingcheng, et al.
Published: (2025)
by: Ni, Jingcheng, et al.
Published: (2025)
Similar Items
-
Open-Pose 3D Zero-Shot Learning: Benchmark and Challenges
by: Zhao, Weiguang, et al.
Published: (2023) -
3D-CDRGP: Towards Cross-Device Robotic Grasping Policy in 3D Open World
by: Zhao, Weiguang, et al.
Published: (2024) -
Consistency Diffusion Models for Single-Image 3D Reconstruction with Priors
by: Jiang, Chenru, et al.
Published: (2025) -
From 2D Images to 3D Model:Weakly Supervised Multi-View Face Reconstruction with Deep Fusion
by: Zhao, Weiguang, et al.
Published: (2022) -
PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly Detection
by: Ye, Jianan, et al.
Published: (2024)