High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Weizhi, Lin, Junfan, Chen, Peixin, Lin, Liang, Li, Guanbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Style-Preserving Lip Sync via Audio-Aware Style Reference
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
von: Wu, Linzhi, et al.
Veröffentlicht: (2024)
von: Wu, Linzhi, et al.
Veröffentlicht: (2024)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face Synthesis
von: Bi, Chongke, et al.
Veröffentlicht: (2024)
von: Bi, Chongke, et al.
Veröffentlicht: (2024)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022)
von: Tang, Anni, et al.
Veröffentlicht: (2022)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
Efficient Low-Resolution Face Recognition via Bridge Distillation
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
von: Liu, Ke, et al.
Veröffentlicht: (2026)
von: Liu, Ke, et al.
Veröffentlicht: (2026)
Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head Generation
von: Tan, Jintao, et al.
Veröffentlicht: (2024)
von: Tan, Jintao, et al.
Veröffentlicht: (2024)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
ReCorD: Reasoning and Correcting Diffusion for HOI Generation
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
2DGS-Avatar: Animatable High-fidelity Clothed Avatar via 2D Gaussian Splatting
von: Yan, Qipeng, et al.
Veröffentlicht: (2025)
von: Yan, Qipeng, et al.
Veröffentlicht: (2025)
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
von: Cho, Kyusun, et al.
Veröffentlicht: (2024)
von: Cho, Kyusun, et al.
Veröffentlicht: (2024)
Palmprint De-Identification Using Diffusion Model for High-Quality and Diverse Synthesis
von: Yan, Licheng, et al.
Veröffentlicht: (2025)
von: Yan, Licheng, et al.
Veröffentlicht: (2025)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
von: Fang, Jianwu, et al.
Veröffentlicht: (2025)
von: Fang, Jianwu, et al.
Veröffentlicht: (2025)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
von: Li, Linfei, et al.
Veröffentlicht: (2025)
von: Li, Linfei, et al.
Veröffentlicht: (2025)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
von: Kang, Fang, et al.
Veröffentlicht: (2025)
von: Kang, Fang, et al.
Veröffentlicht: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
Towards Real-world Video Face Restoration: A New Benchmark
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
Band-Attention Modulated RetNet for Face Forgery Detection
von: Zhang, Zhida, et al.
Veröffentlicht: (2024)
von: Zhang, Zhida, et al.
Veröffentlicht: (2024)
StyleLipSync: Style-based Personalized Lip-sync Video Generation
von: Ki, Taekyung, et al.
Veröffentlicht: (2023)
von: Ki, Taekyung, et al.
Veröffentlicht: (2023)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
CIV-DG: Conditional Instrumental Variables for Domain Generalization in Medical Imaging
von: Bai, Shaojin, et al.
Veröffentlicht: (2026)
von: Bai, Shaojin, et al.
Veröffentlicht: (2026)
From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2025)
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Style-Preserving Lip Sync via Audio-Aware Style Reference
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024) -
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
von: Wu, Linzhi, et al.
Veröffentlicht: (2024) -
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023) -
NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face Synthesis
von: Bi, Chongke, et al.
Veröffentlicht: (2024) -
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
von: Hao, Bowen, et al.
Veröffentlicht: (2025)