Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hong, Fa-Ting, Xu, Zunnan, Zhou, Zixiang, Zhou, Jun, Li, Xiu, Lin, Qin, Lu, Qinglin, Xu, Dan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning Online Scale Transformation for Talking Head Video Generation
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
di: Xu, Zunnan, et al.
Pubblicazione: (2025)
di: Xu, Zunnan, et al.
Pubblicazione: (2025)
Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
di: Zhao, Shuling, et al.
Pubblicazione: (2024)
di: Zhao, Shuling, et al.
Pubblicazione: (2024)
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
di: Xu, Zunnan, et al.
Pubblicazione: (2024)
di: Xu, Zunnan, et al.
Pubblicazione: (2024)
FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model
di: Zhou, Jun, et al.
Pubblicazione: (2025)
di: Zhou, Jun, et al.
Pubblicazione: (2025)
Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos
di: Takahashi, Riku, et al.
Pubblicazione: (2025)
di: Takahashi, Riku, et al.
Pubblicazione: (2025)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
di: Chen, Yi, et al.
Pubblicazione: (2025)
di: Chen, Yi, et al.
Pubblicazione: (2025)
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
di: Wang, Haotian, et al.
Pubblicazione: (2024)
di: Wang, Haotian, et al.
Pubblicazione: (2024)
ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
di: Peng, Ziqiao, et al.
Pubblicazione: (2025)
di: Peng, Ziqiao, et al.
Pubblicazione: (2025)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
di: Zhang, Youliang, et al.
Pubblicazione: (2026)
di: Zhang, Youliang, et al.
Pubblicazione: (2026)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
di: Zhang, Guozhen, et al.
Pubblicazione: (2025)
di: Zhang, Guozhen, et al.
Pubblicazione: (2025)
Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
di: Liu, Zihao, et al.
Pubblicazione: (2025)
di: Liu, Zihao, et al.
Pubblicazione: (2025)
READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
di: Wang, Haotian, et al.
Pubblicazione: (2025)
di: Wang, Haotian, et al.
Pubblicazione: (2025)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
di: Flynn, John, et al.
Pubblicazione: (2026)
di: Flynn, John, et al.
Pubblicazione: (2026)
Dual Audio-Centric Modality Coupling for Talking Head Generation
di: Fu, Ao, et al.
Pubblicazione: (2025)
di: Fu, Ao, et al.
Pubblicazione: (2025)
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
di: Lian, Jiesong, et al.
Pubblicazione: (2026)
di: Lian, Jiesong, et al.
Pubblicazione: (2026)
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
di: Zhang, Ruicheng, et al.
Pubblicazione: (2026)
di: Zhang, Ruicheng, et al.
Pubblicazione: (2026)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
di: Li, Tianqi, et al.
Pubblicazione: (2024)
di: Li, Tianqi, et al.
Pubblicazione: (2024)
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
di: Wang, Yucheng, et al.
Pubblicazione: (2025)
di: Wang, Yucheng, et al.
Pubblicazione: (2025)
ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search
di: Liu, Zhenjie, et al.
Pubblicazione: (2025)
di: Liu, Zhenjie, et al.
Pubblicazione: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
di: Ling, Jun, et al.
Pubblicazione: (2024)
di: Ling, Jun, et al.
Pubblicazione: (2024)
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
di: Fei, Zhengcong, et al.
Pubblicazione: (2025)
di: Fei, Zhengcong, et al.
Pubblicazione: (2025)
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
di: Zhang, Ruicheng, et al.
Pubblicazione: (2025)
di: Zhang, Ruicheng, et al.
Pubblicazione: (2025)
Video Motion Graphs
di: Liu, Haiyang, et al.
Pubblicazione: (2025)
di: Liu, Haiyang, et al.
Pubblicazione: (2025)
EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
di: Liu, Chang, et al.
Pubblicazione: (2025)
di: Liu, Chang, et al.
Pubblicazione: (2025)
USV: Unified Sparsification for Accelerating Video Diffusion Models
di: Wu, Xinjian, et al.
Pubblicazione: (2025)
di: Wu, Xinjian, et al.
Pubblicazione: (2025)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
di: Chopin, Baptiste, et al.
Pubblicazione: (2025)
di: Chopin, Baptiste, et al.
Pubblicazione: (2025)
Identity-Consistent Video Generation under Large Facial-Angle Variations
di: Hu, Bin, et al.
Pubblicazione: (2026)
di: Hu, Bin, et al.
Pubblicazione: (2026)
Euphonium: Steering Video Flow Matching via Process Reward Gradient Guided Stochastic Dynamics
di: Zhong, Ruizhe, et al.
Pubblicazione: (2026)
di: Zhong, Ruizhe, et al.
Pubblicazione: (2026)
BATON: Aligning Text-to-Audio Model with Human Preference Feedback
di: Liao, Huan, et al.
Pubblicazione: (2024)
di: Liao, Huan, et al.
Pubblicazione: (2024)
FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models
di: Aneja, Shivangi, et al.
Pubblicazione: (2023)
di: Aneja, Shivangi, et al.
Pubblicazione: (2023)
Diffusion Models as Masked Audio-Video Learners
di: Nunez, Elvis, et al.
Pubblicazione: (2023)
di: Nunez, Elvis, et al.
Pubblicazione: (2023)
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
di: Xu, Zhihua, et al.
Pubblicazione: (2025)
di: Xu, Zhihua, et al.
Pubblicazione: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
di: Lin, Yukang, et al.
Pubblicazione: (2023)
di: Lin, Yukang, et al.
Pubblicazione: (2023)
Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
di: Xu, Jingyi, et al.
Pubblicazione: (2024)
di: Xu, Jingyi, et al.
Pubblicazione: (2024)
Free-viewpoint Human Animation with Pose-correlated Reference Selection
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos
di: Zhang, Weixia, et al.
Pubblicazione: (2024)
di: Zhang, Weixia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning Online Scale Transformation for Talking Head Video Generation
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024) -
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024) -
HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
di: Xu, Zunnan, et al.
Pubblicazione: (2025) -
Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
di: Zhao, Shuling, et al.
Pubblicazione: (2024) -
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
di: Xu, Zunnan, et al.
Pubblicazione: (2024)