An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Su, Tianhui, Tan, Tien-Ping, Mdhaffar, Salima, Estève, Yannick, Sini, Aghilas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
di: Mdhaffar, Salima, et al.
Pubblicazione: (2024)
di: Mdhaffar, Salima, et al.
Pubblicazione: (2024)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
di: Guo, Zhao, et al.
Pubblicazione: (2025)
di: Guo, Zhao, et al.
Pubblicazione: (2025)
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
di: Zhou, Xuanru, et al.
Pubblicazione: (2024)
di: Zhou, Xuanru, et al.
Pubblicazione: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
di: Zhou, Junzuo, et al.
Pubblicazione: (2024)
di: Zhou, Junzuo, et al.
Pubblicazione: (2024)
MuCodec: Ultra Low-Bitrate Music Codec
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
di: Xin, Detai, et al.
Pubblicazione: (2024)
di: Xin, Detai, et al.
Pubblicazione: (2024)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
CUSIDE-array: A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations
di: Kong, Xiangzhu, et al.
Pubblicazione: (2024)
di: Kong, Xiangzhu, et al.
Pubblicazione: (2024)
MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
di: Li, Jingyu, et al.
Pubblicazione: (2025)
di: Li, Jingyu, et al.
Pubblicazione: (2025)
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
di: Yang, Leyan, et al.
Pubblicazione: (2026)
di: Yang, Leyan, et al.
Pubblicazione: (2026)
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
di: Sekkat, Chloé, et al.
Pubblicazione: (2024)
di: Sekkat, Chloé, et al.
Pubblicazione: (2024)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
di: Wan, Genshun, et al.
Pubblicazione: (2026)
di: Wan, Genshun, et al.
Pubblicazione: (2026)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
Ultra-Low Latency Speech Enhancement - A Comprehensive Study
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
End-to-End Speech Recognition with Pre-trained Masked Language Model
di: Higuchi, Yosuke, et al.
Pubblicazione: (2024)
di: Higuchi, Yosuke, et al.
Pubblicazione: (2024)
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
di: Ren, Yanzhou, et al.
Pubblicazione: (2026)
di: Ren, Yanzhou, et al.
Pubblicazione: (2026)
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding
di: Zhao, Mingyu, et al.
Pubblicazione: (2026)
di: Zhao, Mingyu, et al.
Pubblicazione: (2026)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
di: Guo, Dake, et al.
Pubblicazione: (2025)
di: Guo, Dake, et al.
Pubblicazione: (2025)
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
di: Wang, Pu, et al.
Pubblicazione: (2024)
di: Wang, Pu, et al.
Pubblicazione: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
di: Jing, Kangqi, et al.
Pubblicazione: (2025)
di: Jing, Kangqi, et al.
Pubblicazione: (2025)
Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2022)
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2022)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
di: Sahipjohn, Neha, et al.
Pubblicazione: (2024)
di: Sahipjohn, Neha, et al.
Pubblicazione: (2024)
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
di: Pothula, Aishwarya, et al.
Pubblicazione: (2025)
di: Pothula, Aishwarya, et al.
Pubblicazione: (2025)
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
di: Chen, Peijie, et al.
Pubblicazione: (2025)
di: Chen, Peijie, et al.
Pubblicazione: (2025)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
di: Kong, Xiangzhu, et al.
Pubblicazione: (2025)
di: Kong, Xiangzhu, et al.
Pubblicazione: (2025)
TS3-Codec: Transformer-Based Simple Streaming Single Codec
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
di: Hirano, Yuta, et al.
Pubblicazione: (2025)
di: Hirano, Yuta, et al.
Pubblicazione: (2025)
Sound Source Separation Using Latent Variational Block-Wise Disentanglement
di: Helwani, Karim, et al.
Pubblicazione: (2024)
di: Helwani, Karim, et al.
Pubblicazione: (2024)
CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder
di: Cui, Jianwei, et al.
Pubblicazione: (2024)
di: Cui, Jianwei, et al.
Pubblicazione: (2024)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
di: Li, Hanzhao, et al.
Pubblicazione: (2024)
di: Li, Hanzhao, et al.
Pubblicazione: (2024)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
di: Liu, Fei, et al.
Pubblicazione: (2024)
di: Liu, Fei, et al.
Pubblicazione: (2024)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
di: Moslem, Yasmin
Pubblicazione: (2024)
di: Moslem, Yasmin
Pubblicazione: (2024)
Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
di: Dai, Wang, et al.
Pubblicazione: (2024)
di: Dai, Wang, et al.
Pubblicazione: (2024)
Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2024)
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
di: Mdhaffar, Salima, et al.
Pubblicazione: (2024) -
Decoder-only Architecture for Streaming End-to-end Speech Recognition
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024) -
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
di: Guo, Zhao, et al.
Pubblicazione: (2025) -
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
di: Zhou, Xuanru, et al.
Pubblicazione: (2024) -
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
di: Zhou, Junzuo, et al.
Pubblicazione: (2024)