Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xu, Tang, Shengeng, Wang, Fei, Cheng, Lechao, Guo, Dan, Xue, Feng, Hong, Richang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025)
by: Hao, Bowen, et al.
Published: (2025)
Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
by: Park, Se Jin, et al.
Published: (2023)
by: Park, Se Jin, et al.
Published: (2023)
SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observation
by: Tang, Shengeng, et al.
Published: (2024)
by: Tang, Shengeng, et al.
Published: (2024)
SyncAnyone: Implicit Disentanglement via Progressive Self-Correction for Lip-Syncing in the wild
by: Zhang, Xindi, et al.
Published: (2025)
by: Zhang, Xindi, et al.
Published: (2025)
Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance Disentanglement
by: Yu, Runyi, et al.
Published: (2024)
by: Yu, Runyi, et al.
Published: (2024)
BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
StyleLipSync: Style-based Personalized Lip-sync Video Generation
by: Ki, Taekyung, et al.
Published: (2023)
by: Ki, Taekyung, et al.
Published: (2023)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
by: Huang, Yihuan, et al.
Published: (2025)
by: Huang, Yihuan, et al.
Published: (2025)
Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes
by: Liu, Weifeng, et al.
Published: (2024)
by: Liu, Weifeng, et al.
Published: (2024)
Text-Driven Diffusion Model for Sign Language Production
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
by: He, Jiashu, et al.
Published: (2025)
by: He, Jiashu, et al.
Published: (2025)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
by: Li, Chunyu, et al.
Published: (2024)
by: Li, Chunyu, et al.
Published: (2024)
3DXTalker: Unifying Identity, Lip Sync, Emotion, and Spatial Dynamics in Expressive 3D Talking Avatars
by: Wang, Zhongju, et al.
Published: (2026)
by: Wang, Zhongju, et al.
Published: (2026)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
by: Ling, Zeyu, et al.
Published: (2025)
by: Ling, Zeyu, et al.
Published: (2025)
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
Style-Preserving Lip Sync via Audio-Aware Style Reference
by: Zhong, Weizhi, et al.
Published: (2024)
by: Zhong, Weizhi, et al.
Published: (2024)
GenSync: A Generalized Talking Head Framework for Audio-driven Multi-Subject Lip-Sync using 3D Gaussian Splatting
by: Agarwal, Anushka, et al.
Published: (2025)
by: Agarwal, Anushka, et al.
Published: (2025)
JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
by: Park, Sungjoon, et al.
Published: (2025)
by: Park, Sungjoon, et al.
Published: (2025)
DynamicLip: Shape-Independent Continuous Authentication via Lip Articulator Dynamics
by: Chen, Huashan, et al.
Published: (2025)
by: Chen, Huashan, et al.
Published: (2025)
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025)
by: Zinonos, Andreas, et al.
Published: (2025)
HighSync: High-Quality Lip Synchronization via Latent Diffusion Models
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
by: Wang, Xu, et al.
Published: (2024)
by: Wang, Xu, et al.
Published: (2024)
Talking Face Generation With Lip and Identity Priors
by: Jiajie Wu, et al.
Published: (2025)
by: Jiajie Wu, et al.
Published: (2025)
UniSync: Towards Generalizable and High-Fidelity Lip Synchronization for Challenging Scenarios
by: Fan, Ruidi, et al.
Published: (2026)
by: Fan, Ruidi, et al.
Published: (2026)
Removing Averaging: Personalized Lip-Sync Driven Characters Based on Identity Adapter
by: Zhu, Yanyu, et al.
Published: (2025)
by: Zhu, Yanyu, et al.
Published: (2025)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
by: Kang, Fang, et al.
Published: (2025)
by: Kang, Fang, et al.
Published: (2025)
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
by: Zhong, Weizhi, et al.
Published: (2024)
by: Zhong, Weizhi, et al.
Published: (2024)
SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
by: Yee, Phyo Thet, et al.
Published: (2025)
by: Yee, Phyo Thet, et al.
Published: (2025)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
Modality Alignment Meets Federated Broadcasting
by: Ma, Yuting, et al.
Published: (2024)
by: Ma, Yuting, et al.
Published: (2024)
Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
by: Guo, Mingce, et al.
Published: (2024)
by: Guo, Mingce, et al.
Published: (2024)
Detecting Lip-Syncing Deepfakes: Vision Temporal Transformer for Analyzing Mouth Inconsistencies
by: Datta, Soumyya Kanti, et al.
Published: (2025)
by: Datta, Soumyya Kanti, et al.
Published: (2025)
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
by: Tang, Shengeng, et al.
Published: (2024)
by: Tang, Shengeng, et al.
Published: (2024)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
by: Bigata, Antoni, et al.
Published: (2025)
by: Bigata, Antoni, et al.
Published: (2025)
Similar Items
-
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025) -
Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
by: Park, Se Jin, et al.
Published: (2023) -
SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation
by: Wang, Xu, et al.
Published: (2025) -
CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
by: Wang, Xu, et al.
Published: (2026) -
Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observation
by: Tang, Shengeng, et al.
Published: (2024)