Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhihua, Chen, Tianshui, Yang, Zhijing, Peng, Siyuan, Wang, Keze, Lin, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
by: Lu, Zhenxuan, et al.
Published: (2026)
by: Lu, Zhenxuan, et al.
Published: (2026)
SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation
by: Liu, Yujian, et al.
Published: (2025)
by: Liu, Yujian, et al.
Published: (2025)
Monocular and Generalizable Gaussian Talking Head Animation
by: Gong, Shengjie, et al.
Published: (2025)
by: Gong, Shengjie, et al.
Published: (2025)
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026)
by: Chen, Tianshui, et al.
Published: (2026)
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
by: Zhang, Zeren, et al.
Published: (2024)
by: Zhang, Zeren, et al.
Published: (2024)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
by: Wang, Baiqin, et al.
Published: (2025)
by: Wang, Baiqin, et al.
Published: (2025)
Audio-Synchronized Visual Animation
by: Zhang, Lin, et al.
Published: (2024)
by: Zhang, Lin, et al.
Published: (2024)
Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos
by: Takahashi, Riku, et al.
Published: (2025)
by: Takahashi, Riku, et al.
Published: (2025)
GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
by: Agarwal, Madhav, et al.
Published: (2025)
by: Agarwal, Madhav, et al.
Published: (2025)
FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
by: Wang, MengChao, et al.
Published: (2025)
by: Wang, MengChao, et al.
Published: (2025)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
by: Chen, Shunian, et al.
Published: (2025)
by: Chen, Shunian, et al.
Published: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
by: Ling, Jun, et al.
Published: (2024)
by: Ling, Jun, et al.
Published: (2024)
OT-Talk: Animating 3D Talking Head with Optimal Transportation
by: Wang, Xinmu, et al.
Published: (2025)
by: Wang, Xinmu, et al.
Published: (2025)
Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels
by: Ruan, Haoxian, et al.
Published: (2024)
by: Ruan, Haoxian, et al.
Published: (2024)
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
by: Xu, Mingwang, et al.
Published: (2024)
by: Xu, Mingwang, et al.
Published: (2024)
Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding
by: Zhang, Yuhui, et al.
Published: (2026)
by: Zhang, Yuhui, et al.
Published: (2026)
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
by: Liang, Susan, et al.
Published: (2024)
by: Liang, Susan, et al.
Published: (2024)
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
by: Zhang, Wenli, et al.
Published: (2026)
by: Zhang, Wenli, et al.
Published: (2026)
SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting
by: Wu, Hefeng, et al.
Published: (2023)
by: Wu, Hefeng, et al.
Published: (2023)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024)
by: Nazarieh, Fatemeh, et al.
Published: (2024)
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
by: Vo, Hoang-Son, et al.
Published: (2025)
by: Vo, Hoang-Son, et al.
Published: (2025)
UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control
by: Sun, Wenzhang, et al.
Published: (2024)
by: Sun, Wenzhang, et al.
Published: (2024)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
by: Liu, Xiangyu, et al.
Published: (2026)
by: Liu, Xiangyu, et al.
Published: (2026)
Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026)
by: Chen, Tianshui, et al.
Published: (2026)
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
by: Zhen, Dingcheng, et al.
Published: (2025)
by: Zhen, Dingcheng, et al.
Published: (2025)
Dynamic Correlation Learning and Regularization for Multi-Label Confidence Calibration
by: Chen, Tianshui, et al.
Published: (2024)
by: Chen, Tianshui, et al.
Published: (2024)
OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking
by: Wang, Zhongjian, et al.
Published: (2025)
by: Wang, Zhongjian, et al.
Published: (2025)
TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens
by: Zhao, Qingcheng, et al.
Published: (2026)
by: Zhao, Qingcheng, et al.
Published: (2026)
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation
by: Cao, Xuyang, et al.
Published: (2024)
by: Cao, Xuyang, et al.
Published: (2024)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos
by: Zhang, Weixia, et al.
Published: (2024)
by: Zhang, Weixia, et al.
Published: (2024)
MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation
by: Meng, Dechao, et al.
Published: (2025)
by: Meng, Dechao, et al.
Published: (2025)
Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis
by: Shen, Shuai, et al.
Published: (2025)
by: Shen, Shuai, et al.
Published: (2025)
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
by: Hong, Fa-Ting, et al.
Published: (2025)
by: Hong, Fa-Ting, et al.
Published: (2025)
EmbedTalk: Triplane-Free Talking Head Synthesis using Embedding-Driven Gaussian Deformation
by: Saggar, Arpita, et al.
Published: (2026)
by: Saggar, Arpita, et al.
Published: (2026)
Similar Items
-
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
by: Lu, Zhenxuan, et al.
Published: (2026) -
SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation
by: Liu, Yujian, et al.
Published: (2025) -
Monocular and Generalizable Gaussian Talking Head Animation
by: Gong, Shengjie, et al.
Published: (2025) -
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026) -
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
by: Zhang, Zeren, et al.
Published: (2024)