AUHead: Realistic Emotional Talking Head Generation via Action Units Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Lyu, Jiayi, Qu, Leigang, Zhang, Wenjing, Jiang, Hanyu, Liu, Kai, Zhou, Zhenglin, Xia, Xiaobo, Xue, Jian, Chua, Tat-Seng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
Towards Unified Facial Action Unit Recognition Framework by Large Language Models
di: Hu, Guohong, et al.
Pubblicazione: (2024)
di: Hu, Guohong, et al.
Pubblicazione: (2024)
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
di: Chen, Yiyang, et al.
Pubblicazione: (2022)
di: Chen, Yiyang, et al.
Pubblicazione: (2022)
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
Discriminative Probing and Tuning for Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis
di: Wang, Wenqing, et al.
Pubblicazione: (2025)
di: Wang, Wenqing, et al.
Pubblicazione: (2025)
Principled Multimodal Representation Learning
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
Continual Multimodal Contrastive Learning
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
di: Yang, Wenhao, et al.
Pubblicazione: (2026)
di: Yang, Wenhao, et al.
Pubblicazione: (2026)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
di: Shi, Enyi, et al.
Pubblicazione: (2026)
di: Shi, Enyi, et al.
Pubblicazione: (2026)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
di: Gao, Haowen, et al.
Pubblicazione: (2025)
di: Gao, Haowen, et al.
Pubblicazione: (2025)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
di: He, Jinghan, et al.
Pubblicazione: (2024)
di: He, Jinghan, et al.
Pubblicazione: (2024)
Towards Modality Generalization: A Benchmark and Prospective Analysis
di: Liu, Xiaohao, et al.
Pubblicazione: (2024)
di: Liu, Xiaohao, et al.
Pubblicazione: (2024)
EMOdiffhead: Continuously Emotional Control in Talking Head Generation via Diffusion
di: Zhang, Jian, et al.
Pubblicazione: (2024)
di: Zhang, Jian, et al.
Pubblicazione: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
di: Jin, Zhe, et al.
Pubblicazione: (2025)
di: Jin, Zhe, et al.
Pubblicazione: (2025)
ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
di: Yang, Tianyu, et al.
Pubblicazione: (2026)
di: Yang, Tianyu, et al.
Pubblicazione: (2026)
VINCIE: Unlocking In-context Image Editing from Video
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
di: Lan, Xing, et al.
Pubblicazione: (2024)
di: Lan, Xing, et al.
Pubblicazione: (2024)
EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters
di: Shen, Xuli, et al.
Pubblicazione: (2025)
di: Shen, Xuli, et al.
Pubblicazione: (2025)
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
di: Shi, Hanlei, et al.
Pubblicazione: (2025)
di: Shi, Hanlei, et al.
Pubblicazione: (2025)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
di: Shen, Fei, et al.
Pubblicazione: (2025)
di: Shen, Fei, et al.
Pubblicazione: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
di: Chu, Meng, et al.
Pubblicazione: (2025)
di: Chu, Meng, et al.
Pubblicazione: (2025)
Universal Scene Graph Generation
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
Learning to Ask Critical Questions for Assisting Product Search
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
Calibrated Multimodal Representation Learning with Missing Modalities
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis
di: Cha, Junuk, et al.
Pubblicazione: (2025)
di: Cha, Junuk, et al.
Pubblicazione: (2025)
FreeTalk: Emotional Topology-Free 3D Talking Heads
di: Nocentini, Federico, et al.
Pubblicazione: (2026)
di: Nocentini, Federico, et al.
Pubblicazione: (2026)
EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
di: Liu, Chang, et al.
Pubblicazione: (2025)
di: Liu, Chang, et al.
Pubblicazione: (2025)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
di: Deng, Yang, et al.
Pubblicazione: (2023)
di: Deng, Yang, et al.
Pubblicazione: (2023)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
di: Zhang, An, et al.
Pubblicazione: (2024)
di: Zhang, An, et al.
Pubblicazione: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
di: Zhou, Yuchen, et al.
Pubblicazione: (2025)
di: Zhou, Yuchen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025) -
Towards Unified Facial Action Unit Recognition Framework by Large Language Models
di: Hu, Guohong, et al.
Pubblicazione: (2024) -
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025) -
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025) -
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
di: Chen, Yiyang, et al.
Pubblicazione: (2022)