FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junjie, Hu, Yao, Li, Junjie, Li, Kangyue, Liu, Kun, Li, Wenpeng, Li, Xu, Li, Ziyuan, Shen, Feiyu, Tang, Xu, Wei, Manzhen, Wu, Yichen, Xie, Fenglong, Xu, Kaituo, Xie, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
by: Xie, Kun, et al.
Published: (2025)
by: Xie, Kun, et al.
Published: (2025)
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
by: Xu, Kaituo, et al.
Published: (2026)
by: Xu, Kaituo, et al.
Published: (2026)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
by: Guo, Hao-Han, et al.
Published: (2024)
by: Guo, Hao-Han, et al.
Published: (2024)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
by: Guo, Hao-Han, et al.
Published: (2025)
by: Guo, Hao-Han, et al.
Published: (2025)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
by: Li, Junjie, et al.
Published: (2023)
by: Li, Junjie, et al.
Published: (2023)
$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
by: Ray, Soham, et al.
Published: (2026)
by: Ray, Soham, et al.
Published: (2026)
DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
by: Wei, Haojie, et al.
Published: (2024)
by: Wei, Haojie, et al.
Published: (2024)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
by: Xu, Kai-Tuo, et al.
Published: (2025)
by: Xu, Kai-Tuo, et al.
Published: (2025)
EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge
by: Huang, Hengyan, et al.
Published: (2026)
by: Huang, Hengyan, et al.
Published: (2026)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
by: Guo, Yiwei, et al.
Published: (2024)
by: Guo, Yiwei, et al.
Published: (2024)
U3-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
by: Li, Haiyun, et al.
Published: (2025)
by: Li, Haiyun, et al.
Published: (2025)
UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction
by: Li, Yadong, et al.
Published: (2026)
by: Li, Yadong, et al.
Published: (2026)
Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
by: Wu, Weijie, et al.
Published: (2025)
by: Wu, Weijie, et al.
Published: (2025)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
by: Guo, Haohan, et al.
Published: (2024)
by: Guo, Haohan, et al.
Published: (2024)
FabasedVC: Enhancing Voice Conversion with Text Modality Fusion and Phoneme-Level SSL Features
by: Wang, Wenyu, et al.
Published: (2025)
by: Wang, Wenyu, et al.
Published: (2025)
AVENet: Disentangling Features by Approximating Average Features for Voice Conversion
by: Wang, Wenyu, et al.
Published: (2025)
by: Wang, Wenyu, et al.
Published: (2025)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
by: Sha, Binzhu, et al.
Published: (2023)
by: Sha, Binzhu, et al.
Published: (2023)
Adversarial multi-task underwater acoustic target recognition: towards robustness against various influential factors
by: Xie, Yuan, et al.
Published: (2024)
by: Xie, Yuan, et al.
Published: (2024)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
by: Likhomanenko, Tatiana, et al.
Published: (2025)
by: Likhomanenko, Tatiana, et al.
Published: (2025)
YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis
by: Zhou, Zhitong, et al.
Published: (2025)
by: Zhou, Zhitong, et al.
Published: (2025)
HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
by: Mu, Bingshen, et al.
Published: (2024)
by: Mu, Bingshen, et al.
Published: (2024)
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
by: Yao, Yiqun, et al.
Published: (2025)
by: Yao, Yiqun, et al.
Published: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
by: Xu, Rixi, et al.
Published: (2026)
by: Xu, Rixi, et al.
Published: (2026)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
by: Xu, Haiying, et al.
Published: (2025)
by: Xu, Haiying, et al.
Published: (2025)
CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
by: Hira, Medha, et al.
Published: (2024)
by: Hira, Medha, et al.
Published: (2024)
Controlling your Attributes in Voice
by: Li, Xuyuan, et al.
Published: (2025)
by: Li, Xuyuan, et al.
Published: (2025)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
by: Li, Yuke, et al.
Published: (2024)
by: Li, Yuke, et al.
Published: (2024)
Advancing Robust Underwater Acoustic Target Recognition through Multi-task Learning and Multi-Gate Mixture-of-Experts
by: Xie, Yuan, et al.
Published: (2024)
by: Xie, Yuan, et al.
Published: (2024)
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
by: Hao, Chunbo, et al.
Published: (2026)
by: Hao, Chunbo, et al.
Published: (2026)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
by: Zhou, Li, et al.
Published: (2026)
by: Zhou, Li, et al.
Published: (2026)
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
by: Cui, Wenqian, et al.
Published: (2025)
by: Cui, Wenqian, et al.
Published: (2025)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
by: Tao, Ruijie, et al.
Published: (2024)
by: Tao, Ruijie, et al.
Published: (2024)
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
by: Li, Junjie, et al.
Published: (2025)
by: Li, Junjie, et al.
Published: (2025)
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
by: Li, Xiquan, et al.
Published: (2026)
by: Li, Xiquan, et al.
Published: (2026)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
by: Riera, Pablo, et al.
Published: (2026)
by: Riera, Pablo, et al.
Published: (2026)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
by: Du, Zongyang, et al.
Published: (2024)
by: Du, Zongyang, et al.
Published: (2024)
Marco-Voice Technical Report
by: Tian, Fengping, et al.
Published: (2025)
by: Tian, Fengping, et al.
Published: (2025)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
by: Cao, Junjie, et al.
Published: (2025)
by: Cao, Junjie, et al.
Published: (2025)
Similar Items
-
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
by: Xie, Kun, et al.
Published: (2025) -
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
by: Xu, Kaituo, et al.
Published: (2026) -
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
by: Guo, Hao-Han, et al.
Published: (2024) -
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
by: Guo, Hao-Han, et al.
Published: (2025) -
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
by: Li, Junjie, et al.
Published: (2023)