Learning Online Scale Transformation for Talking Head Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Fa-Ting, Xu, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
by: Zhao, Shuling, et al.
Published: (2024)
by: Zhao, Shuling, et al.
Published: (2024)
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
by: Wang, Haotian, et al.
Published: (2024)
by: Wang, Haotian, et al.
Published: (2024)
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
by: Hong, Fa-Ting, et al.
Published: (2025)
by: Hong, Fa-Ting, et al.
Published: (2025)
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
by: Cheng, Hanbo, et al.
Published: (2024)
by: Cheng, Hanbo, et al.
Published: (2024)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
by: Feng, Guanwen, et al.
Published: (2025)
by: Feng, Guanwen, et al.
Published: (2025)
Passive Dementia Screening via Facial Temporal Micro-Dynamics Analysis of In-the-Wild Talking-Head Video
by: Cenacchi, Filippo, et al.
Published: (2025)
by: Cenacchi, Filippo, et al.
Published: (2025)
Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head Generation
by: Tan, Jintao, et al.
Published: (2024)
by: Tan, Jintao, et al.
Published: (2024)
NLDF: Neural Light Dynamic Fields for Efficient 3D Talking Head Generation
by: Guanchen, Niu
Published: (2024)
by: Guanchen, Niu
Published: (2024)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
by: Ling, Jun, et al.
Published: (2024)
by: Ling, Jun, et al.
Published: (2024)
Adaptive Super Resolution For One-Shot Talking-Head Generation
by: Song, Luchuan, et al.
Published: (2024)
by: Song, Luchuan, et al.
Published: (2024)
Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity
by: Tian, Jiahao, et al.
Published: (2026)
by: Tian, Jiahao, et al.
Published: (2026)
GvT: A Graph-based Vision Transformer with Talking-Heads Utilizing Sparsity, Trained from Scratch on Small Datasets
by: Shan, Dongjing, et al.
Published: (2024)
by: Shan, Dongjing, et al.
Published: (2024)
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
by: Shi, Hanlei, et al.
Published: (2025)
by: Shi, Hanlei, et al.
Published: (2025)
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
by: Wang, Ni, et al.
Published: (2024)
by: Wang, Ni, et al.
Published: (2024)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
by: Wang, Baiqin, et al.
Published: (2025)
by: Wang, Baiqin, et al.
Published: (2025)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
by: Long, Jianzhi, et al.
Published: (2025)
by: Long, Jianzhi, et al.
Published: (2025)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
by: Menapace, Willi, et al.
Published: (2024)
by: Menapace, Willi, et al.
Published: (2024)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
by: Marmon, Andrew, et al.
Published: (2024)
by: Marmon, Andrew, et al.
Published: (2024)
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
by: Wang, Yucheng, et al.
Published: (2025)
by: Wang, Yucheng, et al.
Published: (2025)
Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2025)
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2025)
Controllable Talking Face Generation by Implicit Facial Keypoints Editing
by: Zhao, Dong, et al.
Published: (2024)
by: Zhao, Dong, et al.
Published: (2024)
SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis
by: Jia, Peng, et al.
Published: (2026)
by: Jia, Peng, et al.
Published: (2026)
Video-T1: Test-Time Scaling for Video Generation
by: Liu, Fangfu, et al.
Published: (2025)
by: Liu, Fangfu, et al.
Published: (2025)
Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video Conference
by: Li, Jianglong, et al.
Published: (2026)
by: Li, Jianglong, et al.
Published: (2026)
MAGI-1: Autoregressive Video Generation at Scale
by: ai, Sand., et al.
Published: (2025)
by: ai, Sand., et al.
Published: (2025)
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
by: Zhang, Zeren, et al.
Published: (2024)
by: Zhang, Zeren, et al.
Published: (2024)
DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Multi-Head Explainer: A General Framework to Improve Explainability in CNNs and Transformers
by: Sun, Bohang, et al.
Published: (2025)
by: Sun, Bohang, et al.
Published: (2025)
Watch and Learn: Learning to Use Computers from Online Videos
by: Song, Chan Hee, et al.
Published: (2025)
by: Song, Chan Hee, et al.
Published: (2025)
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
by: Chen, Hejia, et al.
Published: (2025)
by: Chen, Hejia, et al.
Published: (2025)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
Text-Driven Emotionally Continuous Talking Face Generation
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
by: Weng, Yuzhe, et al.
Published: (2026)
by: Weng, Yuzhe, et al.
Published: (2026)
Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy Prediction
by: Yan, Chi, et al.
Published: (2025)
by: Yan, Chi, et al.
Published: (2025)
Similar Items
-
Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
by: Zhao, Shuling, et al.
Published: (2024) -
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
by: Wang, Haotian, et al.
Published: (2024) -
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
by: Hong, Fa-Ting, et al.
Published: (2025) -
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
by: Cheng, Hanbo, et al.
Published: (2024) -
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
by: Xie, Yifan, et al.
Published: (2025)