DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Ziqiao, Fan, Yanbo, Wu, Haoyu, Wang, Xuan, Liu, Hongyan, He, Jun, Fan, Zhaoxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
Dual Audio-Centric Modality Coupling for Talking Head Generation
von: Fu, Ao, et al.
Veröffentlicht: (2025)
von: Fu, Ao, et al.
Veröffentlicht: (2025)
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
von: Meng, Ming, et al.
Veröffentlicht: (2025)
von: Meng, Ming, et al.
Veröffentlicht: (2025)
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
von: Yang, Kaixing, et al.
Veröffentlicht: (2024)
von: Yang, Kaixing, et al.
Veröffentlicht: (2024)
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
von: Xue, Ke, et al.
Veröffentlicht: (2025)
von: Xue, Ke, et al.
Veröffentlicht: (2025)
MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
von: Wang, Haotian, et al.
Veröffentlicht: (2025)
von: Wang, Haotian, et al.
Veröffentlicht: (2025)
From Independence to Interaction: Speaker-Aware Simulation of Multi-Speaker Conversational Timing
von: Gedeon, Máté, et al.
Veröffentlicht: (2025)
von: Gedeon, Máté, et al.
Veröffentlicht: (2025)
Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis
von: Shen, Shuai, et al.
Veröffentlicht: (2025)
von: Shen, Shuai, et al.
Veröffentlicht: (2025)
Cross-Talk Reduction
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2024)
Warm Chat: Diffuse Emotion-aware Interactive Talking Head Avatar with Tree-Structured Guidance
von: Yang, Haijie, et al.
Veröffentlicht: (2025)
von: Yang, Haijie, et al.
Veröffentlicht: (2025)
Cross-Talk Speech Reduction, by Separation, for Separation
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2026)
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2026)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
von: Yu, Fan, et al.
Veröffentlicht: (2025)
von: Yu, Fan, et al.
Veröffentlicht: (2025)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
von: Kang, Fang, et al.
Veröffentlicht: (2025)
von: Kang, Fang, et al.
Veröffentlicht: (2025)
Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
von: Zhang, Bingyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Bingyuan, et al.
Veröffentlicht: (2024)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart Speakers
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
von: Liu, Bei, et al.
Veröffentlicht: (2024)
von: Liu, Bei, et al.
Veröffentlicht: (2024)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Residual Speaker Representation for One-Shot Voice Conversion
von: Xu, Le, et al.
Veröffentlicht: (2023)
von: Xu, Le, et al.
Veröffentlicht: (2023)
ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
TalkPlay-Tools: Conversational Music Recommendation with LLM Tool Calling
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
von: Wen, Wen, et al.
Veröffentlicht: (2024)
von: Wen, Wen, et al.
Veröffentlicht: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
von: Xue, Ke, et al.
Veröffentlicht: (2026)
von: Xue, Ke, et al.
Veröffentlicht: (2026)
AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
von: Fu, Linya, et al.
Veröffentlicht: (2025)
von: Fu, Linya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
von: Zhou, Xukun, et al.
Veröffentlicht: (2024) -
Dual Audio-Centric Modality Coupling for Talking Head Generation
von: Fu, Ao, et al.
Veröffentlicht: (2025) -
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
von: Meng, Ming, et al.
Veröffentlicht: (2025) -
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
von: Yang, Kaixing, et al.
Veröffentlicht: (2025) -
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025)