AI killed the video star. Audio-driven diffusion model for expressive talking head generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chopin, Baptiste, Dhamija, Tashvik, Balaji, Pranav, Wang, Yaohui, Dantcheva, Antitza |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
by: Bora, Maheswar, et al.
Published: (2025)
by: Bora, Maheswar, et al.
Published: (2025)
THEval. Evaluation Framework for Talking Head Video Generation
by: Quignon, Nabyl, et al.
Published: (2025)
by: Quignon, Nabyl, et al.
Published: (2025)
Beyond Real versus Fake Towards Intent-Aware Video Analysis
by: Atreya, Saurabh, et al.
Published: (2025)
by: Atreya, Saurabh, et al.
Published: (2025)
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
by: Egin, Anil, et al.
Published: (2026)
by: Egin, Anil, et al.
Published: (2026)
LIA-X: Interpretable Latent Portrait Animator
by: Wang, Yaohui, et al.
Published: (2025)
by: Wang, Yaohui, et al.
Published: (2025)
LEO: Generative Latent Image Animator for Human Video Synthesis
by: Wang, Yaohui, et al.
Published: (2023)
by: Wang, Yaohui, et al.
Published: (2023)
LAC: Latent Action Composition for Skeleton-based Action Segmentation
by: Yang, Di, et al.
Published: (2023)
by: Yang, Di, et al.
Published: (2023)
Beyond the Visible: A Survey on Cross-spectral Face Recognition
by: Anghelone, David, et al.
Published: (2022)
by: Anghelone, David, et al.
Published: (2022)
DenVisCoM: Dense Vision Correspondence Mamba for Efficient and Real-time Optical Flow and Stereo Estimation
by: Anand, Tushar, et al.
Published: (2026)
by: Anand, Tushar, et al.
Published: (2026)
AM Flow: Adapters for Temporal Processing in Action Recognition
by: Agrawal, Tanay, et al.
Published: (2024)
by: Agrawal, Tanay, et al.
Published: (2024)
HFNeRF: Learning Human Biomechanic Features with Neural Radiance Fields
by: Dey, Arnab, et al.
Published: (2024)
by: Dey, Arnab, et al.
Published: (2024)
GHNeRF: Learning Generalizable Human Features with Efficient Neural Radiance Fields
by: Dey, Arnab, et al.
Published: (2024)
by: Dey, Arnab, et al.
Published: (2024)
ReactionMamba: Generating Short & Long Human Reaction Sequences
by: Beg, Hajra Anwar, et al.
Published: (2025)
by: Beg, Hajra Anwar, et al.
Published: (2025)
Towards motion from video diffusion models
by: Janson, Paul, et al.
Published: (2024)
by: Janson, Paul, et al.
Published: (2024)
Just Dance with $π$! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection
by: Majhi, Snehashis, et al.
Published: (2025)
by: Majhi, Snehashis, et al.
Published: (2025)
MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
by: Ye, Zhenhui, et al.
Published: (2024)
by: Ye, Zhenhui, et al.
Published: (2024)
Performance of Gaussian Mixture Model Classifiers on Embedded Feature Spaces
by: Chopin, Jeremy, et al.
Published: (2024)
by: Chopin, Jeremy, et al.
Published: (2024)
Spintronics for image recognition: performance benchmarking via ultrafast data-driven simulations
by: Moureaux, Anatole, et al.
Published: (2023)
by: Moureaux, Anatole, et al.
Published: (2023)
Towards multi-modal forgery representation learning for AI-generated video detection and localization
by: Le, Dat, et al.
Published: (2026)
by: Le, Dat, et al.
Published: (2026)
Dynamic watermarks in images generated by diffusion models
by: Chen, Yunzhuo, et al.
Published: (2025)
by: Chen, Yunzhuo, et al.
Published: (2025)
Emotion recognition in talking-face videos using persistent entropy and neural networks
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2021)
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2021)
MVOC: a training-free multiple video object composition method with diffusion models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Autoregression-free video prediction using diffusion model for mitigating error propagation
by: Ko, Woonho, et al.
Published: (2025)
by: Ko, Woonho, et al.
Published: (2025)
Taming generative video models for zero-shot optical flow extraction
by: Kim, Seungwoo, et al.
Published: (2025)
by: Kim, Seungwoo, et al.
Published: (2025)
Generative diffusion models for agricultural AI: plant image generation, indoor-to-outdoor translation, and expert preference alignment
by: Tan, Da, et al.
Published: (2025)
by: Tan, Da, et al.
Published: (2025)
Learned representation-guided diffusion models for large-image generation
by: Graikos, Alexandros, et al.
Published: (2023)
by: Graikos, Alexandros, et al.
Published: (2023)
Dark Miner: Defend against undesirable generation for text-to-image diffusion models
by: Meng, Zheling, et al.
Published: (2024)
by: Meng, Zheling, et al.
Published: (2024)
From Macro to Micro: Boosting micro-expression recognition via pre-training on macro-expression videos
by: Li, Hanting, et al.
Published: (2024)
by: Li, Hanting, et al.
Published: (2024)
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
Boosting Audio-visual Zero-shot Learning with Large Language Models
by: Chen, Haoxing, et al.
Published: (2023)
by: Chen, Haoxing, et al.
Published: (2023)
Restereo: Diffusion stereo video generation and restoration
by: Huang, Xingchang, et al.
Published: (2025)
by: Huang, Xingchang, et al.
Published: (2025)
Can video generation replace cinematographers? Research on the cinematic language of generated video
by: Li, Xiaozhe, et al.
Published: (2024)
by: Li, Xiaozhe, et al.
Published: (2024)
GHOST 2.0: generative high-fidelity one shot transfer of heads
by: Groshev, Alexander, et al.
Published: (2025)
by: Groshev, Alexander, et al.
Published: (2025)
Flow caching for autoregressive video generation
by: Ma, Yuexiao, et al.
Published: (2026)
by: Ma, Yuexiao, et al.
Published: (2026)
Comprehensive Review of EEG-to-Output Research: Decoding Neural Signals into Images, Videos, and Audio
by: Sabharwal, Yashvir, et al.
Published: (2024)
by: Sabharwal, Yashvir, et al.
Published: (2024)
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
by: Li, Shujia, et al.
Published: (2025)
by: Li, Shujia, et al.
Published: (2025)
Multimodal generative semantic communication based on latent diffusion model
by: Fu, Weiqi, et al.
Published: (2024)
by: Fu, Weiqi, et al.
Published: (2024)
AI driven shadow model detection in agropv farms
by: Dornadula, Sai Paavan Kumar, et al.
Published: (2023)
by: Dornadula, Sai Paavan Kumar, et al.
Published: (2023)
Non-uniform Point Cloud Upsampling via Local Manifold Distribution
by: Fang, Yaohui, et al.
Published: (2025)
by: Fang, Yaohui, et al.
Published: (2025)
Similar Items
-
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025) -
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
by: Bora, Maheswar, et al.
Published: (2025) -
THEval. Evaluation Framework for Talking Head Video Generation
by: Quignon, Nabyl, et al.
Published: (2025) -
Beyond Real versus Fake Towards Intent-Aware Video Analysis
by: Atreya, Saurabh, et al.
Published: (2025) -
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
by: Egin, Anil, et al.
Published: (2026)