SyncDiff: Diffusion-based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Xulin, Gao, Heting, Chen, Ziyi, Chang, Peng, Han, Mei, Hasegawa-Johnson, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
by: He, Wenkun, et al.
Published: (2024)
by: He, Wenkun, et al.
Published: (2024)
SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
by: Peng, Ziqiao, et al.
Published: (2023)
by: Peng, Ziqiao, et al.
Published: (2023)
SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
by: Mazumdar, Soumya, et al.
Published: (2026)
by: Mazumdar, Soumya, et al.
Published: (2026)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
by: Fan, Xulin, et al.
Published: (2026)
by: Fan, Xulin, et al.
Published: (2026)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
by: Huang, Yihuan, et al.
Published: (2025)
by: Huang, Yihuan, et al.
Published: (2025)
EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
SyncSDE: A Probabilistic Framework for Diffusion Synchronization
by: Lee, Hyunjun, et al.
Published: (2025)
by: Lee, Hyunjun, et al.
Published: (2025)
InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
by: Hong, Fa-Ting, et al.
Published: (2024)
by: Hong, Fa-Ting, et al.
Published: (2024)
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
Towards Unsupervised Speech Recognition Without Pronunciation Models
by: Ni, Junrui, et al.
Published: (2024)
by: Ni, Junrui, et al.
Published: (2024)
UniSync: A Unified Framework for Audio-Visual Synchronization
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
RocSync: Millisecond-Accurate Temporal Synchronization for Heterogeneous Camera Systems
by: Meyer, Jaro, et al.
Published: (2025)
by: Meyer, Jaro, et al.
Published: (2025)
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
by: Vo, Hoang-Son, et al.
Published: (2025)
by: Vo, Hoang-Son, et al.
Published: (2025)
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis
by: Gao, Huan-ang, et al.
Published: (2024)
by: Gao, Huan-ang, et al.
Published: (2024)
SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis
by: Jia, Peng, et al.
Published: (2026)
by: Jia, Peng, et al.
Published: (2026)
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
by: Zhang, Wenli, et al.
Published: (2026)
by: Zhang, Wenli, et al.
Published: (2026)
ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search
by: Liu, Zhenjie, et al.
Published: (2025)
by: Liu, Zhenjie, et al.
Published: (2025)
GenSync: A Generalized Talking Head Framework for Audio-driven Multi-Subject Lip-Sync using 3D Gaussian Splatting
by: Agarwal, Anushka, et al.
Published: (2025)
by: Agarwal, Anushka, et al.
Published: (2025)
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces
by: Yeo, Kyeongmin, et al.
Published: (2025)
by: Yeo, Kyeongmin, et al.
Published: (2025)
SyncTweedies: A General Generative Framework Based on Synchronized Diffusions
by: Kim, Jaihoon, et al.
Published: (2024)
by: Kim, Jaihoon, et al.
Published: (2024)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
by: Cong, Gaoxiang, et al.
Published: (2026)
by: Cong, Gaoxiang, et al.
Published: (2026)
JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
by: Park, Sungjoon, et al.
Published: (2025)
by: Park, Sungjoon, et al.
Published: (2025)
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
by: Xu, Zhongweiyang, et al.
Published: (2025)
by: Xu, Zhongweiyang, et al.
Published: (2025)
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
by: Xu, Zhihua, et al.
Published: (2025)
by: Xu, Zhihua, et al.
Published: (2025)
DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
by: Liu, Shaowei, et al.
Published: (2025)
by: Liu, Shaowei, et al.
Published: (2025)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
by: Ling, Zeyu, et al.
Published: (2025)
by: Ling, Zeyu, et al.
Published: (2025)
HighSync: High-Quality Lip Synchronization via Latent Diffusion Models
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
by: Qian, Kaizhi, et al.
Published: (2025)
by: Qian, Kaizhi, et al.
Published: (2025)
SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
SyncVIS: Synchronized Video Instance Segmentation
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
by: Li, Chunyu, et al.
Published: (2024)
by: Li, Chunyu, et al.
Published: (2024)
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
OT-Talk: Animating 3D Talking Head with Optimal Transportation
by: Wang, Xinmu, et al.
Published: (2025)
by: Wang, Xinmu, et al.
Published: (2025)
Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head Generation
by: Tan, Jintao, et al.
Published: (2024)
by: Tan, Jintao, et al.
Published: (2024)
Similar Items
-
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
by: He, Wenkun, et al.
Published: (2024) -
SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
by: Peng, Ziqiao, et al.
Published: (2023) -
SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
by: Peng, Ziqiao, et al.
Published: (2025) -
TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
by: Mazumdar, Soumya, et al.
Published: (2026) -
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
by: Fan, Xulin, et al.
Published: (2026)