Recent Advances in Discrete Speech Tokens: A Review
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Yiwei, Li, Zhihan, Wang, Hankun, Li, Bohan, Shao, Chongtian, Zhang, Hanglei, Du, Chenpeng, Chen, Xie, Liu, Shujie, Yu, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
von: Wang, Hankun, et al.
Veröffentlicht: (2025)
von: Wang, Hankun, et al.
Veröffentlicht: (2025)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
von: Li, Bohan, et al.
Veröffentlicht: (2025)
von: Li, Bohan, et al.
Veröffentlicht: (2025)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2025)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2025)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
von: Li, Jizhen, et al.
Veröffentlicht: (2024)
von: Li, Jizhen, et al.
Veröffentlicht: (2024)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
von: Zhuang, Xuanyu, et al.
Veröffentlicht: (2024)
von: Zhuang, Xuanyu, et al.
Veröffentlicht: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
von: Premananth, Gowtham, et al.
Veröffentlicht: (2023)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2023)
Conformer-based Ultrasound-to-Speech Conversion
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
SonicVisionLM: Playing Sound with Vision Language Models
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
Building Audio-Visual Digital Twins with Smartphones
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
von: Wang, Hankun, et al.
Veröffentlicht: (2025) -
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024) -
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024) -
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023) -
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)