KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
Fuente:
arXiv
Saved in:
| Main Authors: | Bigata, Antoni, Stypułkowski, Michał, Mira, Rodrigo, Bounareli, Stella, Vougioukas, Konstantinos, Landgraf, Zoe, Drobyshev, Nikita, Zieba, Maciej, Petridis, Stavros, Pantic, Maja |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
by: Bigata, Antoni, et al.
Published: (2025)
by: Bigata, Antoni, et al.
Published: (2025)
EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
by: Drobyshev, Nikita, et al.
Published: (2024)
by: Drobyshev, Nikita, et al.
Published: (2024)
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025)
by: Zinonos, Andreas, et al.
Published: (2025)
FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion
by: Mishima, Kazuaki, et al.
Published: (2025)
by: Mishima, Kazuaki, et al.
Published: (2025)
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
by: Chen, Honglie, et al.
Published: (2024)
by: Chen, Honglie, et al.
Published: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026)
by: Cappellazzo, Umberto, et al.
Published: (2026)
Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
by: Seo, Junyoung, et al.
Published: (2025)
by: Seo, Junyoung, et al.
Published: (2025)
Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
by: Anand, et al.
Published: (2025)
by: Anand, et al.
Published: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
AutoLoRA: AutoGuidance Meets Low-Rank Adaptation for Diffusion Models
by: Kasymov, Artur, et al.
Published: (2024)
by: Kasymov, Artur, et al.
Published: (2024)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
Fast Heuristic Scheduling and Trajectory Planning for Robotic Fruit Harvesters with Multiple Cartesian Arms
by: Zhu, Yuankai, et al.
Published: (2025)
by: Zhu, Yuankai, et al.
Published: (2025)
Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2026)
by: Haliassos, Alexandros, et al.
Published: (2026)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Expressive Speech-driven Facial Animation with controllable emotions
by: Chen, Yutong, et al.
Published: (2023)
by: Chen, Yutong, et al.
Published: (2023)
Neural Distance-Guided Path Integral Control for Tractor-Trailer Navigation
by: Wei, Peng, et al.
Published: (2026)
by: Wei, Peng, et al.
Published: (2026)
KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding
by: Xu, Zhihao, et al.
Published: (2024)
by: Xu, Zhihao, et al.
Published: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
by: Cappellazzo, Umberto, et al.
Published: (2024)
by: Cappellazzo, Umberto, et al.
Published: (2024)
Corporate Author Entry Records Retrieved by Use of Derived Truncated Search Keys
by: Landgraf, Alan L., et al.
Published: (1973)
by: Landgraf, Alan L., et al.
Published: (1973)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
TreeFlow: Going beyond Tree-based Gaussian Probabilistic Regression
by: Wielopolski, Patryk, et al.
Published: (2022)
by: Wielopolski, Patryk, et al.
Published: (2022)
Sextactic and type-9 points on the Fermat cubic and associated objects
by: Merta, Łukasz, et al.
Published: (2024)
by: Merta, Łukasz, et al.
Published: (2024)
Catalog Records Retrieved by Personal Author Using Derived Search Keys
by: Landgraf, Alan L., et al.
Published: (1973)
by: Landgraf, Alan L., et al.
Published: (1973)
TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens
by: Zhao, Qingcheng, et al.
Published: (2026)
by: Zhao, Qingcheng, et al.
Published: (2026)
Vision-based Navigation of Unmanned Aerial Vehicles in Orchards: An Imitation Learning Approach
by: Wei, Peng, et al.
Published: (2025)
by: Wei, Peng, et al.
Published: (2025)
Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Large-scale unsupervised audio pre-training for video-to-speech synthesis
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
One-shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models
by: Li, Sibo, et al.
Published: (2025)
by: Li, Sibo, et al.
Published: (2025)
Towards plausibility in time series counterfactual explanations
by: Kostrzewa, Marcin, et al.
Published: (2026)
by: Kostrzewa, Marcin, et al.
Published: (2026)
A Probabilistic Consensus-Driven Approach for Robust Counterfactual Explanations
by: Kostrzewa, Marcin, et al.
Published: (2026)
by: Kostrzewa, Marcin, et al.
Published: (2026)
Counterfactual Explanations Under Concept Drift
by: Kostrzewa, Marcin, et al.
Published: (2026)
by: Kostrzewa, Marcin, et al.
Published: (2026)
Dynamic Data Pruning for Automatic Speech Recognition
by: Xiao, Qiao, et al.
Published: (2024)
by: Xiao, Qiao, et al.
Published: (2024)
Chapter From the Lab to the Real World: Affect Recognition Using Multiple Cues and Modalities
by: Gunes, Hatice, et al.
Published: (2021)
by: Gunes, Hatice, et al.
Published: (2021)
Similar Items
-
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
by: Bigata, Antoni, et al.
Published: (2025) -
EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
by: Drobyshev, Nikita, et al.
Published: (2024) -
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025) -
FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion
by: Mishima, Kazuaki, et al.
Published: (2025) -
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
by: Haliassos, Alexandros, et al.
Published: (2024)