Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Xiang, Pang, Youxin, Zhao, Xiaochen, Xu, Chao, Wang, Lizhen, Xiao, Hongjiang, Yan, Shi, Zhang, Hongwen, Liu, Yebin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GMTalker: Gaussian Mixture-based Audio-Driven Emotional Talking Video Portraits
by: Xia, Yibo, et al.
Published: (2023)
by: Xia, Yibo, et al.
Published: (2023)
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
by: Chen, Yushuo, et al.
Published: (2025)
by: Chen, Yushuo, et al.
Published: (2025)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
by: Pang, Youxin, et al.
Published: (2024)
by: Pang, Youxin, et al.
Published: (2024)
Recovering 3D Human Mesh from Monocular Images: A Survey
by: Tian, Yating, et al.
Published: (2022)
by: Tian, Yating, et al.
Published: (2022)
StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
InvertAvatar: Incremental GAN Inversion for Generalized Head Avatars
by: Zhao, Xiaochen, et al.
Published: (2023)
by: Zhao, Xiaochen, et al.
Published: (2023)
GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction
by: Xu, Chao, et al.
Published: (2026)
by: Xu, Chao, et al.
Published: (2026)
GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
by: Guo, Yujie, et al.
Published: (2025)
by: Guo, Yujie, et al.
Published: (2025)
Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing
by: Shi, Hao, et al.
Published: (2026)
by: Shi, Hao, et al.
Published: (2026)
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
by: Deng, Xiang, et al.
Published: (2026)
by: Deng, Xiang, et al.
Published: (2026)
FPED: A Functional-Network Prior-Guided Mixture-of-Experts Framework for Interpretable Brain Decoding
by: Ren, Yudan, et al.
Published: (2026)
by: Ren, Yudan, et al.
Published: (2026)
Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation
by: Lin, Siyou, et al.
Published: (2026)
by: Lin, Siyou, et al.
Published: (2026)
ManiDext: Hand-Object Manipulation Synthesis via Continuous Correspondence Embeddings and Residual-Guided Diffusion
by: Zhang, Jiajun, et al.
Published: (2024)
by: Zhang, Jiajun, et al.
Published: (2024)
Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers
by: Mittal, Manan, et al.
Published: (2025)
by: Mittal, Manan, et al.
Published: (2025)
W-HMR: Monocular Human Mesh Recovery in World Space with Weak-Supervised Calibration
by: Yao, Wei, et al.
Published: (2023)
by: Yao, Wei, et al.
Published: (2023)
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
Domain-Expert-Guided Hybrid Mixture-of-Experts for Medical AI: Integrating Data-Driven Learning with Clinical Priors
by: Gu, Jinchen, et al.
Published: (2026)
by: Gu, Jinchen, et al.
Published: (2026)
UniTalker: Conversational Speech-Visual Synthesis
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
by: Li, Ronghui, et al.
Published: (2024)
by: Li, Ronghui, et al.
Published: (2024)
Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
AGMA: Adaptive Gaussian Mixture Anchors for Prior-Guided Multimodal Human Trajectory Forecasting
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
by: Min, Dongchan, et al.
Published: (2022)
by: Min, Dongchan, et al.
Published: (2022)
MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
by: Cheng, Junda, et al.
Published: (2025)
by: Cheng, Junda, et al.
Published: (2025)
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
by: Cho, Kyusun, et al.
Published: (2024)
by: Cho, Kyusun, et al.
Published: (2024)
UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model
by: Fan, Xiangyu, et al.
Published: (2024)
by: Fan, Xiangyu, et al.
Published: (2024)
SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields
by: Li, Qijing, et al.
Published: (2025)
by: Li, Qijing, et al.
Published: (2025)
Ins-HOI: Instance Aware Human-Object Interactions Recovery
by: Zhang, Jiajun, et al.
Published: (2023)
by: Zhang, Jiajun, et al.
Published: (2023)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)
by: Chai, Ying, et al.
Published: (2025)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
Animatable and Relightable Gaussians for High-fidelity Human Avatar Modeling
by: Li, Zhe, et al.
Published: (2023)
by: Li, Zhe, et al.
Published: (2023)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
by: Berghi, Davide, et al.
Published: (2024)
by: Berghi, Davide, et al.
Published: (2024)
MoEScore: Mixture-of-Experts-Based Text-Audio Relevance Score Prediction for Text-to-Audio System Evaluation
by: Sun, Bochao, et al.
Published: (2026)
by: Sun, Bochao, et al.
Published: (2026)
Audio-Guided Dynamic Modality Fusion with Stereo-Aware Attention for Audio-Visual Navigation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis
by: Guo, Yuhang, et al.
Published: (2025)
by: Guo, Yuhang, et al.
Published: (2025)
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
by: Zhou, Hao, et al.
Published: (2024)
by: Zhou, Hao, et al.
Published: (2024)
Similar Items
-
GMTalker: Gaussian Mixture-based Audio-Driven Emotional Talking Video Portraits
by: Xia, Yibo, et al.
Published: (2023) -
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
by: Chen, Yushuo, et al.
Published: (2025) -
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024) -
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025) -
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025)