StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yifan, Cen, Zhi, Peng, Sida, Chen, Xiangwei, Deng, Yifu, Zhu, Xinyu, Jia, Fan, Zhou, Xiaowei, Bao, Hujun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation
by: Hu, Yongzhen, et al.
Published: (2025)
by: Hu, Yongzhen, et al.
Published: (2025)
Generating Human Motion in 3D Scenes from Text Descriptions
by: Cen, Zhi, et al.
Published: (2024)
by: Cen, Zhi, et al.
Published: (2024)
GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer
by: Lin, Yihong, et al.
Published: (2024)
by: Lin, Yihong, et al.
Published: (2024)
UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model
by: Fan, Xiangyu, et al.
Published: (2024)
by: Fan, Xiangyu, et al.
Published: (2024)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation
by: Cen, Zhi, et al.
Published: (2025)
by: Cen, Zhi, et al.
Published: (2025)
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
by: Jin, Yudong, et al.
Published: (2025)
by: Jin, Yudong, et al.
Published: (2025)
World-Grounded Human Motion Recovery via Gravity-View Coordinates
by: Shen, Zehong, et al.
Published: (2024)
by: Shen, Zehong, et al.
Published: (2024)
Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space
by: Xiao, Lixing, et al.
Published: (2025)
by: Xiao, Lixing, et al.
Published: (2025)
SAiD: Speech-driven Blendshape Facial Animation with Diffusion
by: Park, Inkyu, et al.
Published: (2023)
by: Park, Inkyu, et al.
Published: (2023)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
by: Ye, Zhen, et al.
Published: (2026)
by: Ye, Zhen, et al.
Published: (2026)
MaPa: Text-driven Photorealistic Material Painting for 3D Shapes
by: Zhang, Shangzan, et al.
Published: (2024)
by: Zhang, Shangzan, et al.
Published: (2024)
HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
by: Yu, Zhiyuan, et al.
Published: (2025)
by: Yu, Zhiyuan, et al.
Published: (2025)
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
by: Liu, Tao, et al.
Published: (2024)
by: Liu, Tao, et al.
Published: (2024)
Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
by: Deng, Xiang, et al.
Published: (2024)
by: Deng, Xiang, et al.
Published: (2024)
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
by: Zhen, Dingcheng, et al.
Published: (2025)
by: Zhen, Dingcheng, et al.
Published: (2025)
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
by: Xu, Zhen, et al.
Published: (2024)
by: Xu, Zhen, et al.
Published: (2024)
MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
by: He, Xingyi, et al.
Published: (2025)
by: He, Xingyi, et al.
Published: (2025)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
by: Wang, Baiqin, et al.
Published: (2025)
by: Wang, Baiqin, et al.
Published: (2025)
FreeTimeGS: Free Gaussian Primitives at Anytime and Anywhere for Dynamic Scene Reconstruction
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Contact Matrix: Enhancing Dance Motion Synthesis with Precise Interaction Modeling
by: Chen, Xuhai, et al.
Published: (2026)
by: Chen, Xuhai, et al.
Published: (2026)
RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
by: Du, Fangyu, et al.
Published: (2025)
by: Du, Fangyu, et al.
Published: (2025)
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
by: Zhang, Yixuan, et al.
Published: (2025)
by: Zhang, Yixuan, et al.
Published: (2025)
Content and Style Aware Audio-Driven Facial Animation
by: Liu, Qingju, et al.
Published: (2024)
by: Liu, Qingju, et al.
Published: (2024)
Precise Action-to-Video Generation Through Visual Action Prompts
by: Wang, Yuang, et al.
Published: (2025)
by: Wang, Yuang, et al.
Published: (2025)
Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
by: Guo, Haoyu, et al.
Published: (2025)
by: Guo, Haoyu, et al.
Published: (2025)
DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer
by: Ma, Zhiyuan, et al.
Published: (2024)
by: Ma, Zhiyuan, et al.
Published: (2024)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
Dyn-E: Local Appearance Editing of Dynamic Neural Radiance Fields
by: Zhang, Shangzan, et al.
Published: (2023)
by: Zhang, Shangzan, et al.
Published: (2023)
Audio Driven Real-Time Facial Animation for Social Telepresence
by: Lee, Jiye, et al.
Published: (2025)
by: Lee, Jiye, et al.
Published: (2025)
Expressive Speech-driven Facial Animation with controllable emotions
by: Chen, Yutong, et al.
Published: (2023)
by: Chen, Yutong, et al.
Published: (2023)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
by: Lei, Ke, et al.
Published: (2026)
by: Lei, Ke, et al.
Published: (2026)
AnimateMe: 4D Facial Expressions via Diffusion Models
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
EnvGS: Modeling View-Dependent Appearance with Environment Gaussian
by: Xie, Tao, et al.
Published: (2024)
by: Xie, Tao, et al.
Published: (2024)
EmoFace: Audio-driven Emotional 3D Face Animation
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
Similar Items
-
Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation
by: Hu, Yongzhen, et al.
Published: (2025) -
Generating Human Motion in 3D Scenes from Text Descriptions
by: Cen, Zhi, et al.
Published: (2024) -
GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer
by: Lin, Yihong, et al.
Published: (2024) -
UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model
by: Fan, Xiangyu, et al.
Published: (2024) -
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)