FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion
Fuente:
arXiv
Saved in:
| Main Authors: | Mishima, Kazuaki, Casademunt, Antoni Bigata, Petridis, Stavros, Pantic, Maja, Suzuki, Kenji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
by: Drobyshev, Nikita, et al.
Published: (2024)
by: Drobyshev, Nikita, et al.
Published: (2024)
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025)
by: Zinonos, Andreas, et al.
Published: (2025)
KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
by: Bigata, Antoni, et al.
Published: (2025)
by: Bigata, Antoni, et al.
Published: (2025)
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
by: Bigata, Antoni, et al.
Published: (2025)
by: Bigata, Antoni, et al.
Published: (2025)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026)
by: Cappellazzo, Umberto, et al.
Published: (2026)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
by: Anand, et al.
Published: (2025)
by: Anand, et al.
Published: (2025)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
by: Chen, Honglie, et al.
Published: (2024)
by: Chen, Honglie, et al.
Published: (2024)
BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion
by: Chang, Di, et al.
Published: (2023)
by: Chang, Di, et al.
Published: (2023)
FaceCraft4D: Animated 3D Facial Avatar Generation from a Single Image
by: Yin, Fei, et al.
Published: (2025)
by: Yin, Fei, et al.
Published: (2025)
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Hearing Loss Detection from Facial Expressions in One-on-one Conversations
by: Yin, Yufeng, et al.
Published: (2024)
by: Yin, Yufeng, et al.
Published: (2024)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Investigating Identity Signals in Conversational Facial Dynamics via Disentangled Expression Features
by: Chapariniya, Masoumeh, et al.
Published: (2025)
by: Chapariniya, Masoumeh, et al.
Published: (2025)
Auditing Facial Emotion Recognition Datasets for Posed Expressions and Racial Bias
by: Khan, Rina, et al.
Published: (2025)
by: Khan, Rina, et al.
Published: (2025)
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
by: Wang, Yuanzhi, et al.
Published: (2026)
by: Wang, Yuanzhi, et al.
Published: (2026)
Crafter: Facial Feature Crafting against Inversion-based Identity Theft on Deep Models
by: Wang, Shiming, et al.
Published: (2024)
by: Wang, Shiming, et al.
Published: (2024)
DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing
by: Jia, Haozhe, et al.
Published: (2023)
by: Jia, Haozhe, et al.
Published: (2023)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
by: Mao, Qing, et al.
Published: (2025)
by: Mao, Qing, et al.
Published: (2025)
Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose Generation
by: Liang, Jiadong, et al.
Published: (2024)
by: Liang, Jiadong, et al.
Published: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
by: Cappellazzo, Umberto, et al.
Published: (2024)
by: Cappellazzo, Umberto, et al.
Published: (2024)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
by: Zhong, Yong, et al.
Published: (2024)
by: Zhong, Yong, et al.
Published: (2024)
IP-FaceDiff: Identity-Preserving Facial Video Editing with Diffusion
by: Anand, Tharun, et al.
Published: (2025)
by: Anand, Tharun, et al.
Published: (2025)
Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Smile upon the Face but Sadness in the Eyes: Emotion Recognition based on Facial Expressions and Eye Behaviors
by: Liu, Yuanyuan, et al.
Published: (2024)
by: Liu, Yuanyuan, et al.
Published: (2024)
Emotion Separation and Recognition from a Facial Expression by Generating the Poker Face with Vision Transformers
by: Li, Jia, et al.
Published: (2022)
by: Li, Jia, et al.
Published: (2022)
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
by: Tang, Jiapeng, et al.
Published: (2025)
by: Tang, Jiapeng, et al.
Published: (2025)
High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning
by: He, Dailan, et al.
Published: (2025)
by: He, Dailan, et al.
Published: (2025)
Emotion Diffusion Classifier with Adaptive Margin Discrepancy Training for Facial Expression Recognition
by: Dong, Rongkang, et al.
Published: (2026)
by: Dong, Rongkang, et al.
Published: (2026)
MagicFace: High-Fidelity Facial Expression Editing with Action-Unit Control
by: Wei, Mengting, et al.
Published: (2025)
by: Wei, Mengting, et al.
Published: (2025)
Pose and Facial Expression Transfer by using StyleGAN
by: Jahoda, Petr, et al.
Published: (2025)
by: Jahoda, Petr, et al.
Published: (2025)
FaceMixup: Enhancing Facial Expression Recognition through Mixed Face Regularization
by: Faria, Fabio A., et al.
Published: (2024)
by: Faria, Fabio A., et al.
Published: (2024)
Safeguarding Facial Identity against Diffusion-based Face Swapping via Cascading Pathway Disruption
by: Wang, Liqin, et al.
Published: (2026)
by: Wang, Liqin, et al.
Published: (2026)
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
EmojiDiff: Advanced Facial Expression Control with High Identity Preservation in Portrait Generation
by: Jiang, Liangwei, et al.
Published: (2024)
by: Jiang, Liangwei, et al.
Published: (2024)
Similar Items
-
EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
by: Drobyshev, Nikita, et al.
Published: (2024) -
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025) -
KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
by: Bigata, Antoni, et al.
Published: (2025) -
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
by: Bigata, Antoni, et al.
Published: (2025) -
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026)