The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Yiwei, Wang, Chenrun, Yang, Yifan, Wang, Hankun, Ma, Ziyang, Du, Chenpeng, Wang, Shuai, Li, Hanzheng, Fan, Shuai, Zhang, Hui, Chen, Xie, Yu, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
von: Chang, Xuankai, et al.
Veröffentlicht: (2024)
von: Chang, Xuankai, et al.
Veröffentlicht: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
Interspeech 2025 URGENT Speech Enhancement Challenge
von: Saijo, Kohei, et al.
Veröffentlicht: (2025)
von: Saijo, Kohei, et al.
Veröffentlicht: (2025)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Direct Preference Optimization for Speech Autoregressive Diffusion Models
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
Recent Advances in Discrete Speech Tokens: A Review
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
A Survey on Speech Large Language Models for Understanding
von: Peng, Jing, et al.
Veröffentlicht: (2024)
von: Peng, Jing, et al.
Veröffentlicht: (2024)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
von: Wang, Hankun, et al.
Veröffentlicht: (2025)
von: Wang, Hankun, et al.
Veröffentlicht: (2025)
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
von: Xiao, Yunchong, et al.
Veröffentlicht: (2026)
von: Xiao, Yunchong, et al.
Veröffentlicht: (2026)
Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
Data Augmentation for End-to-end Code-switching Speech Recognition
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
The SJTU X-LANCE Lab System for MSR Challenge 2025
von: Zhu, Jinxuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jinxuan, et al.
Veröffentlicht: (2026)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
von: Li, Bohan, et al.
Veröffentlicht: (2025)
von: Li, Bohan, et al.
Veröffentlicht: (2025)
Spoken-Term Discovery using Discrete Speech Units
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
Advances in Speech Separation: Techniques, Challenges, and Future Trends
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
Estimating the Completeness of Discrete Speech Units
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2024)
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2024)
Helsinki Speech Challenge 2024
von: Ludvigsen, Martin, et al.
Veröffentlicht: (2024)
von: Ludvigsen, Martin, et al.
Veröffentlicht: (2024)
AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
von: Poncelet, Jakob, et al.
Veröffentlicht: (2024)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 2025
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
von: Chang, Xuankai, et al.
Veröffentlicht: (2024) -
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024) -
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024) -
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024) -
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)