AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Yuxin, Sun, Jiayang, Zhu, Guibo, Cao, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided Distillation
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
von: Mazumdar, Soumya, et al.
Veröffentlicht: (2026)
von: Mazumdar, Soumya, et al.
Veröffentlicht: (2026)
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026)
von: Li, Haitian, et al.
Veröffentlicht: (2026)
Logit Distillation on Manifolds: Mapping by Learning
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
Enabling Automatic Self-Talk Detection via Earables
von: Lee, Euihyeok, et al.
Veröffentlicht: (2025)
von: Lee, Euihyeok, et al.
Veröffentlicht: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
von: Sienkiewicz, Bruno, et al.
Veröffentlicht: (2026)
von: Sienkiewicz, Bruno, et al.
Veröffentlicht: (2026)
SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing
von: Tan, Jiaye, et al.
Veröffentlicht: (2025)
von: Tan, Jiaye, et al.
Veröffentlicht: (2025)
MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
von: Ristori, Eleonora, et al.
Veröffentlicht: (2025)
von: Ristori, Eleonora, et al.
Veröffentlicht: (2025)
Survey on the Evaluation of Generative Models in Music
von: Lerch, Alexander, et al.
Veröffentlicht: (2025)
von: Lerch, Alexander, et al.
Veröffentlicht: (2025)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
Anchored Cyclic Generation: A Novel Paradigm for Long-Sequence Symbolic Music Generation
von: Cao, Boyu, et al.
Veröffentlicht: (2026)
von: Cao, Boyu, et al.
Veröffentlicht: (2026)
Mixer is more than just a model
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
Warm Chat: Diffuse Emotion-aware Interactive Talking Head Avatar with Tree-Structured Guidance
von: Yang, Haijie, et al.
Veröffentlicht: (2025)
von: Yang, Haijie, et al.
Veröffentlicht: (2025)
VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories
von: Zhang, Qian, et al.
Veröffentlicht: (2026)
von: Zhang, Qian, et al.
Veröffentlicht: (2026)
CoMoSVC: Consistency Model-based Singing Voice Conversion
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
S-SONDO: Self-Supervised Knowledge Distillation for General Audio Foundation Models
von: Adlouni, Mohammed Ali El, et al.
Veröffentlicht: (2026)
von: Adlouni, Mohammed Ali El, et al.
Veröffentlicht: (2026)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2025)
E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2025)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
von: Marincione, Davide, et al.
Veröffentlicht: (2025)
von: Marincione, Davide, et al.
Veröffentlicht: (2025)
Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation
von: Si, Chen, et al.
Veröffentlicht: (2025)
von: Si, Chen, et al.
Veröffentlicht: (2025)
Yin-Yang: Developing Motifs With Long-Term Structure And Controllability
von: Bhandari, Keshav, et al.
Veröffentlicht: (2025)
von: Bhandari, Keshav, et al.
Veröffentlicht: (2025)
Mitigating Unauthorized Speech Synthesis for Voice Protection
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided Distillation
von: Chen, Peng, et al.
Veröffentlicht: (2025) -
TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
von: Mazumdar, Soumya, et al.
Veröffentlicht: (2026) -
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025) -
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024) -
PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)