SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yan, Ruiqi, Chen, Wenxi, Liu, Zhanxun, Ma, Ziyang, Lin, Haopeng, Wen, Hanlin, Xie, Hanke, Wu, Jun, Liang, Yuzhe, Zhao, Yuxiang, Feng, Pengchao, Qian, Jiale, Meng, Hao, Dai, Yuhang, Yin, Shunshun, Tao, Ming, Xie, Lei, Yu, Kai, Wang, Xinsheng, Chen, Xie |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
par: Dai, Yuhang, et autres
Publié: (2026)
par: Dai, Yuhang, et autres
Publié: (2026)
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
par: Xie, Hanke, et autres
Publié: (2025)
par: Xie, Hanke, et autres
Publié: (2025)
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
par: Qian, Jiale, et autres
Publié: (2026)
par: Qian, Jiale, et autres
Publié: (2026)
Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
par: Dai, Yuhang, et autres
Publié: (2026)
par: Dai, Yuhang, et autres
Publié: (2026)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
par: Chen, Wenxi, et autres
Publié: (2025)
par: Chen, Wenxi, et autres
Publié: (2025)
SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads
par: Yu, Tan, et autres
Publié: (2026)
par: Yu, Tan, et autres
Publié: (2026)
SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation
par: Shen, Le, et autres
Publié: (2025)
par: Shen, Le, et autres
Publié: (2025)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
par: Chen, Wenxi, et autres
Publié: (2024)
par: Chen, Wenxi, et autres
Publié: (2024)
SoulX-LiveAct: Towards Hour-Scale Real-Time Human Animation with Neighbor Forcing and ConvKV Memory
par: Zhen, Dingcheng, et autres
Publié: (2026)
par: Zhen, Dingcheng, et autres
Publié: (2026)
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
par: Feng, Pengchao, et autres
Publié: (2025)
par: Feng, Pengchao, et autres
Publié: (2025)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
par: Chen, Wenxi, et autres
Publié: (2024)
par: Chen, Wenxi, et autres
Publié: (2024)
PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Full-Duplex Dialogue
par: Jeon, Hyunbae, et autres
Publié: (2026)
par: Jeon, Hyunbae, et autres
Publié: (2026)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
par: Li, Xiquan, et autres
Publié: (2024)
par: Li, Xiquan, et autres
Publié: (2024)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
par: Chen, Wenxi, et autres
Publié: (2024)
par: Chen, Wenxi, et autres
Publié: (2024)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
par: Yan, Ruiqi, et autres
Publié: (2025)
par: Yan, Ruiqi, et autres
Publié: (2025)
MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
par: Li, Xiquan, et autres
Publié: (2025)
par: Li, Xiquan, et autres
Publié: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
par: Tian, Wenjie, et autres
Publié: (2025)
par: Tian, Wenjie, et autres
Publié: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
par: Chen, Jiaxing, et autres
Publié: (2024)
par: Chen, Jiaxing, et autres
Publié: (2024)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
par: Liu, Xuyang, et autres
Publié: (2025)
par: Liu, Xuyang, et autres
Publié: (2025)
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
par: Zheng, Qixi, et autres
Publié: (2026)
par: Zheng, Qixi, et autres
Publié: (2026)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
par: Liu, Zhanxun, et autres
Publié: (2025)
par: Liu, Zhanxun, et autres
Publié: (2025)
Towards Flow-Matching-based TTS without Classifier-Free Guidance
par: Liang, Yuzhe, et autres
Publié: (2025)
par: Liang, Yuzhe, et autres
Publié: (2025)
Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
par: Li, Xiquan, et autres
Publié: (2026)
par: Li, Xiquan, et autres
Publié: (2026)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
par: Tu, Wenming, et autres
Publié: (2025)
par: Tu, Wenming, et autres
Publié: (2025)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
par: Xie, Zichen, et autres
Publié: (2026)
par: Xie, Zichen, et autres
Publié: (2026)
Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
par: Wu, Weijie, et autres
Publié: (2025)
par: Wu, Weijie, et autres
Publié: (2025)
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
par: Lin, Guan-Ting, et autres
Publié: (2026)
par: Lin, Guan-Ting, et autres
Publié: (2026)
LogLite: Lightweight Plug-and-Play Streaming Log Compression
par: Tang, Benzhao, et autres
Publié: (2025)
par: Tang, Benzhao, et autres
Publié: (2025)
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
par: Xie, Rui, et autres
Publié: (2026)
par: Xie, Rui, et autres
Publié: (2026)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
par: Luo, Katie, et autres
Publié: (2025)
par: Luo, Katie, et autres
Publié: (2025)
STDec: Spatio-Temporal Stability Guided Decoding for dLLMs
par: Chen, Yuzhe, et autres
Publié: (2026)
par: Chen, Yuzhe, et autres
Publié: (2026)
Generation-Augmented Generation: A Plug-and-Play Framework for Private Knowledge Injection in Large Language Models
par: Li, Rongji, et autres
Publié: (2026)
par: Li, Rongji, et autres
Publié: (2026)
Joint Beamforming Design and Resource Allocation for IRS-Assisted Full-Duplex Terahertz Systems
par: Qiu, Chi, et autres
Publié: (2025)
par: Qiu, Chi, et autres
Publié: (2025)
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
par: Li, Xiquan, et autres
Publié: (2026)
par: Li, Xiquan, et autres
Publié: (2026)
SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees
par: Xie, Yi, et autres
Publié: (2026)
par: Xie, Yi, et autres
Publié: (2026)
Histological characterization of rat vocal fold across different postnatal periods
par: Xumao Li, et autres
Publié: (2024)
par: Xumao Li, et autres
Publié: (2024)
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
par: Chen, Junjie, et autres
Publié: (2025)
par: Chen, Junjie, et autres
Publié: (2025)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
par: Deng, Keqi, et autres
Publié: (2025)
par: Deng, Keqi, et autres
Publié: (2025)
Documents similaires
-
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
par: Dai, Yuhang, et autres
Publié: (2026) -
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
par: Xie, Hanke, et autres
Publié: (2025) -
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
par: Qian, Jiale, et autres
Publié: (2026) -
Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
par: Dai, Yuhang, et autres
Publié: (2026) -
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
par: Chen, Wenxi, et autres
Publié: (2025)