The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Guo, Yiwei, Wang, Chenrun, Yang, Yifan, Wang, Hankun, Ma, Ziyang, Du, Chenpeng, Wang, Shuai, Li, Hanzheng, Fan, Shuai, Zhang, Hui, Chen, Xie, Yu, Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
par: Chang, Xuankai, et autres
Publié: (2024)
par: Chang, Xuankai, et autres
Publié: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
par: Du, Chenpeng, et autres
Publié: (2024)
par: Du, Chenpeng, et autres
Publié: (2024)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
par: Nakata, Wataru, et autres
Publié: (2024)
par: Nakata, Wataru, et autres
Publié: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
par: Shen, Feiyu, et autres
Publié: (2023)
par: Shen, Feiyu, et autres
Publié: (2023)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
Interspeech 2025 URGENT Speech Enhancement Challenge
par: Saijo, Kohei, et autres
Publié: (2025)
par: Saijo, Kohei, et autres
Publié: (2025)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
Direct Preference Optimization for Speech Autoregressive Diffusion Models
par: Liu, Zhijun, et autres
Publié: (2025)
par: Liu, Zhijun, et autres
Publié: (2025)
Recent Advances in Discrete Speech Tokens: A Review
par: Guo, Yiwei, et autres
Publié: (2025)
par: Guo, Yiwei, et autres
Publié: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
par: Du, Chenpeng, et autres
Publié: (2022)
par: Du, Chenpeng, et autres
Publié: (2022)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
par: Du, Chenpeng, et autres
Publié: (2021)
par: Du, Chenpeng, et autres
Publié: (2021)
A Survey on Speech Large Language Models for Understanding
par: Peng, Jing, et autres
Publié: (2024)
par: Peng, Jing, et autres
Publié: (2024)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
par: Guo, Yiwei, et autres
Publié: (2023)
par: Guo, Yiwei, et autres
Publié: (2023)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
par: Wang, Haoran, et autres
Publié: (2025)
par: Wang, Haoran, et autres
Publié: (2025)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
par: Wang, Hankun, et autres
Publié: (2025)
par: Wang, Hankun, et autres
Publié: (2025)
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
par: Xiao, Yunchong, et autres
Publié: (2026)
par: Xiao, Yunchong, et autres
Publié: (2026)
Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
Data Augmentation for End-to-end Code-switching Speech Recognition
par: Du, Chenpeng, et autres
Publié: (2020)
par: Du, Chenpeng, et autres
Publié: (2020)
The SJTU X-LANCE Lab System for MSR Challenge 2025
par: Zhu, Jinxuan, et autres
Publié: (2026)
par: Zhu, Jinxuan, et autres
Publié: (2026)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
par: Wang, Yongqi, et autres
Publié: (2023)
par: Wang, Yongqi, et autres
Publié: (2023)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
par: Li, Bohan, et autres
Publié: (2025)
par: Li, Bohan, et autres
Publié: (2025)
Spoken-Term Discovery using Discrete Speech Units
par: van Niekerk, Benjamin, et autres
Publié: (2024)
par: van Niekerk, Benjamin, et autres
Publié: (2024)
Advances in Speech Separation: Techniques, Challenges, and Future Trends
par: Li, Kai, et autres
Publié: (2025)
par: Li, Kai, et autres
Publié: (2025)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
par: Tao, Dehua, et autres
Publié: (2024)
par: Tao, Dehua, et autres
Publié: (2024)
Estimating the Completeness of Discrete Speech Units
par: Yeh, Sung-Lin, et autres
Publié: (2024)
par: Yeh, Sung-Lin, et autres
Publié: (2024)
Helsinki Speech Challenge 2024
par: Ludvigsen, Martin, et autres
Publié: (2024)
par: Ludvigsen, Martin, et autres
Publié: (2024)
AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
par: Guo, Yiwei, et autres
Publié: (2025)
par: Guo, Yiwei, et autres
Publié: (2025)
Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
par: Poncelet, Jakob, et autres
Publié: (2024)
par: Poncelet, Jakob, et autres
Publié: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 2025
par: Peng, Yizhou, et autres
Publié: (2025)
par: Peng, Yizhou, et autres
Publié: (2025)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
par: Yang, Yifan, et autres
Publié: (2024)
par: Yang, Yifan, et autres
Publié: (2024)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
par: Dinkel, Heinrich, et autres
Publié: (2026)
par: Dinkel, Heinrich, et autres
Publié: (2026)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
par: Yang, Yifan, et autres
Publié: (2026)
par: Yang, Yifan, et autres
Publié: (2026)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
par: Chen, Peikun, et autres
Publié: (2024)
par: Chen, Peikun, et autres
Publié: (2024)
Documents similaires
-
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
par: Chang, Xuankai, et autres
Publié: (2024) -
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
par: Wang, Hankun, et autres
Publié: (2024) -
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
par: Du, Chenpeng, et autres
Publié: (2024) -
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024) -
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
par: Nakata, Wataru, et autres
Publié: (2024)