AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Zhizhou, Ji, Yicheng, Kong, Zhe, Liu, Yiying, Wang, Jiarui, Feng, Jiasun, Liu, Lupeng, Wang, Xiangyi, Li, Yanjia, She, Yuqing, Qin, Ying, Li, Huan, Mao, Shuiyang, Liu, Wei, Luo, Wenhan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
by: Ji, Yicheng, et al.
Published: (2026)
by: Ji, Yicheng, et al.
Published: (2026)
Is Generative AI an Existential Threat to Human Creatives? Insights from Financial Economics
by: Li, Jiasun
Published: (2024)
by: Li, Jiasun
Published: (2024)
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
by: Liu, Tao, et al.
Published: (2024)
by: Liu, Tao, et al.
Published: (2024)
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
by: Zhang, Ruicheng, et al.
Published: (2026)
by: Zhang, Ruicheng, et al.
Published: (2026)
From Cheap Geometry to Expensive Physics: Elevating Neural Operators via Latent Shape Pretraining
by: Zhang, Zhizhou, et al.
Published: (2025)
by: Zhang, Zhizhou, et al.
Published: (2025)
Trusted-Execution Environment (TEE) for Solving the Replication Crisis in Academia
by: Li, Jiasun, et al.
Published: (2026)
by: Li, Jiasun, et al.
Published: (2026)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
by: Shen, Yijun, et al.
Published: (2025)
by: Shen, Yijun, et al.
Published: (2025)
HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis
by: Liu, Shiyu, et al.
Published: (2025)
by: Liu, Shiyu, et al.
Published: (2025)
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans
by: Zhou, Yingjie, et al.
Published: (2025)
by: Zhou, Yingjie, et al.
Published: (2025)
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
by: Feng, Guanwen, et al.
Published: (2025)
by: Feng, Guanwen, et al.
Published: (2025)
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
by: Zhang, Bingyuan, et al.
Published: (2024)
by: Zhang, Bingyuan, et al.
Published: (2024)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
by: Xiong, Lingyu, et al.
Published: (2024)
by: Xiong, Lingyu, et al.
Published: (2024)
See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
by: Wang, Jinting, et al.
Published: (2025)
by: Wang, Jinting, et al.
Published: (2025)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
by: Meng, Lingwei, et al.
Published: (2024)
by: Meng, Lingwei, et al.
Published: (2024)
DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided Distillation
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
by: Zhou, Yingjie, et al.
Published: (2025)
by: Zhou, Yingjie, et al.
Published: (2025)
M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
by: Jiang, Kui, et al.
Published: (2025)
by: Jiang, Kui, et al.
Published: (2025)
Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
by: Ge, Shuheng, et al.
Published: (2024)
by: Ge, Shuheng, et al.
Published: (2024)
Numerical Unique Ergodicity of Monotone SDEs driven by Nondegenerate Multiplicative Noise
by: Liu, Zhihui, et al.
Published: (2024)
by: Liu, Zhihui, et al.
Published: (2024)
UniTalker: Conversational Speech-Visual Synthesis
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
by: Ye, Zhen, et al.
Published: (2026)
by: Ye, Zhen, et al.
Published: (2026)
LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
by: Feng, Guanwen, et al.
Published: (2024)
by: Feng, Guanwen, et al.
Published: (2024)
OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models
by: Kong, Zhe, et al.
Published: (2024)
by: Kong, Zhe, et al.
Published: (2024)
Quantum heat transport in nonequilibrium anisotropic Dicke model
by: Junran, Kong, et al.
Published: (2026)
by: Junran, Kong, et al.
Published: (2026)
Constructing Ionic Transport Network via Supramolecular Composite Binder in Cathode for All‐Solid‐State Lithium Batteries
by: Haixing Liu, et al.
Published: (2025)
by: Haixing Liu, et al.
Published: (2025)
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
by: Yang, Shaoshu, et al.
Published: (2025)
by: Yang, Shaoshu, et al.
Published: (2025)
ResumeFlow: An LLM-facilitated Pipeline for Personalized Resume Generation and Refinement
by: Zinjad, Saurabh Bhausaheb, et al.
Published: (2024)
by: Zinjad, Saurabh Bhausaheb, et al.
Published: (2024)
OpenClaw-RL: Train Any Agent Simply by Talking
by: Wang, Yinjie, et al.
Published: (2026)
by: Wang, Yinjie, et al.
Published: (2026)
OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking
by: Wang, Zhongjian, et al.
Published: (2025)
by: Wang, Zhongjian, et al.
Published: (2025)
Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing
by: Shi, Hao, et al.
Published: (2026)
by: Shi, Hao, et al.
Published: (2026)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
by: Du, Chenpeng, et al.
Published: (2023)
by: Du, Chenpeng, et al.
Published: (2023)
Automated Constraint Specification for Job Scheduling by Regulating Generative Model with Domain-Specific Representation
by: Shi, Yu-Zhe, et al.
Published: (2025)
by: Shi, Yu-Zhe, et al.
Published: (2025)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
by: Kong, Quan, et al.
Published: (2026)
by: Kong, Quan, et al.
Published: (2026)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
Track Any Motions under Any Disturbances
by: Zhang, Zhikai, et al.
Published: (2025)
by: Zhang, Zhikai, et al.
Published: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
by: Ling, Jun, et al.
Published: (2024)
by: Ling, Jun, et al.
Published: (2024)
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
by: Lu, Yuxin, et al.
Published: (2026)
by: Lu, Yuxin, et al.
Published: (2026)
Similar Items
-
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
by: Ji, Yicheng, et al.
Published: (2026) -
Is Generative AI an Existential Threat to Human Creatives? Insights from Financial Economics
by: Li, Jiasun
Published: (2024) -
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
by: Liu, Tao, et al.
Published: (2024) -
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
by: Zhang, Ruicheng, et al.
Published: (2026) -
From Cheap Geometry to Expensive Physics: Elevating Neural Operators via Latent Shape Pretraining
by: Zhang, Zhizhou, et al.
Published: (2025)