CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Gaojie, Jiang, Jianwen, Liang, Chao, Zhong, Tianyun, Yang, Jiaqi, Zheng, Yanbo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
by: Jiang, Jianwen, et al.
Published: (2024)
by: Jiang, Jianwen, et al.
Published: (2024)
FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
by: Zhong, Tianyun, et al.
Published: (2024)
by: Zhong, Tianyun, et al.
Published: (2024)
MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices
by: Jiang, Jianwen, et al.
Published: (2024)
by: Jiang, Jianwen, et al.
Published: (2024)
Superior and Pragmatic Talking Face Generation with Teacher-Student Framework
by: Liang, Chao, et al.
Published: (2024)
by: Liang, Chao, et al.
Published: (2024)
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
by: Lin, Gaojie, et al.
Published: (2025)
by: Lin, Gaojie, et al.
Published: (2025)
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
by: Jiang, Jianwen, et al.
Published: (2025)
by: Jiang, Jianwen, et al.
Published: (2025)
R3-Avatar: Record and Retrieve Temporal Codebook for Reconstructing Photorealistic Human Avatars
by: Zhan, Yifan, et al.
Published: (2025)
by: Zhan, Yifan, et al.
Published: (2025)
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
by: Liang, Chao, et al.
Published: (2025)
by: Liang, Chao, et al.
Published: (2025)
UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control
by: Sun, Wenzhang, et al.
Published: (2024)
by: Sun, Wenzhang, et al.
Published: (2024)
Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation
by: Zheng, Bowen, et al.
Published: (2026)
by: Zheng, Bowen, et al.
Published: (2026)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
by: Li, Chunyu, et al.
Published: (2024)
by: Li, Chunyu, et al.
Published: (2024)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
by: Gong, Chaoqun, et al.
Published: (2024)
by: Gong, Chaoqun, et al.
Published: (2024)
OHTA: One-shot Hand Avatar via Data-driven Implicit Priors
by: Zheng, Xiaozheng, et al.
Published: (2024)
by: Zheng, Xiaozheng, et al.
Published: (2024)
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
by: Zheng, Jun, et al.
Published: (2024)
by: Zheng, Jun, et al.
Published: (2024)
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
by: Yu, Haojie, et al.
Published: (2025)
by: Yu, Haojie, et al.
Published: (2025)
3DTeethSAM: Taming SAM2 for 3D Teeth Segmentation
by: Lu, Zhiguo, et al.
Published: (2025)
by: Lu, Zhiguo, et al.
Published: (2025)
Pruning for Robust Concept Erasing in Diffusion Models
by: Yang, Tianyun, et al.
Published: (2024)
by: Yang, Tianyun, et al.
Published: (2024)
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
by: Huang, Yubo, et al.
Published: (2025)
by: Huang, Yubo, et al.
Published: (2025)
MCGA: Mixture of Codebooks Hyperspectral Reconstruction via Grayscale-Aware Attention
by: Yang, Zhanjiang, et al.
Published: (2025)
by: Yang, Zhanjiang, et al.
Published: (2025)
AvatarArtist: Open-Domain 4D Avatarization
by: Liu, Hongyu, et al.
Published: (2025)
by: Liu, Hongyu, et al.
Published: (2025)
InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation
by: Lin, Yukang, et al.
Published: (2025)
by: Lin, Yukang, et al.
Published: (2025)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
by: Zuo, Tongchun, et al.
Published: (2025)
by: Zuo, Tongchun, et al.
Published: (2025)
ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance
by: Yang, Haijie, et al.
Published: (2024)
by: Yang, Haijie, et al.
Published: (2024)
AvatarVTON: 4D Virtual Try-On for Animatable Avatars
by: Jiang, Zicheng, et al.
Published: (2025)
by: Jiang, Zicheng, et al.
Published: (2025)
Audio-Driven Universal Gaussian Head Avatars
by: Teotia, Kartik, et al.
Published: (2025)
by: Teotia, Kartik, et al.
Published: (2025)
Taming Diffusion for Dataset Distillation with High Representativeness
by: Zhao, Lin, et al.
Published: (2025)
by: Zhao, Lin, et al.
Published: (2025)
JoyStreamer-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion
by: Li, Chaochao, et al.
Published: (2025)
by: Li, Chaochao, et al.
Published: (2025)
Taming Flow Matching with Unbalanced Optimal Transport into Fast Pansharpening
by: Cao, Zihan, et al.
Published: (2025)
by: Cao, Zihan, et al.
Published: (2025)
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
by: Song, Yiren, et al.
Published: (2024)
by: Song, Yiren, et al.
Published: (2024)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
DiffusionAvatars: Deferred Diffusion for High-fidelity 3D Head Avatars
by: Kirschstein, Tobias, et al.
Published: (2023)
by: Kirschstein, Tobias, et al.
Published: (2023)
Taming Diffusion Models for Image Restoration: A Review
by: Luo, Ziwei, et al.
Published: (2024)
by: Luo, Ziwei, et al.
Published: (2024)
X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio
by: Zhang, Chenxu, et al.
Published: (2025)
by: Zhang, Chenxu, et al.
Published: (2025)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos
by: Chen, Yushuo, et al.
Published: (2024)
by: Chen, Yushuo, et al.
Published: (2024)
UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
by: Chu, Xuangeng, et al.
Published: (2025)
by: Chu, Xuangeng, et al.
Published: (2025)
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
by: Luo, Xianrui, et al.
Published: (2025)
by: Luo, Xianrui, et al.
Published: (2025)
Similar Items
-
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
by: Jiang, Jianwen, et al.
Published: (2024) -
FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
by: Zhong, Tianyun, et al.
Published: (2024) -
MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices
by: Jiang, Jianwen, et al.
Published: (2024) -
Superior and Pragmatic Talking Face Generation with Teacher-Student Framework
by: Liang, Chao, et al.
Published: (2024) -
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
by: Lin, Gaojie, et al.
Published: (2025)