On-device Streaming Discrete Speech Units
Fuente:
arXiv
Guardado en:
| Autores principales: | Choi, Kwanghee, Someki, Masao, Strubell, Emma, Watanabe, Shinji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Discrete Speech Unit Extraction via Independent Component Analysis
por: Nakamura, Tomohiko, et al.
Publicado: (2025)
por: Nakamura, Tomohiko, et al.
Publicado: (2025)
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
por: Someki, Masao, et al.
Publicado: (2023)
por: Someki, Masao, et al.
Publicado: (2023)
ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
por: Someki, Masao, et al.
Publicado: (2024)
por: Someki, Masao, et al.
Publicado: (2024)
Multi-blank Transducers for Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2022)
por: Xu, Hainan, et al.
Publicado: (2022)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
por: Chang, Xuankai, et al.
Publicado: (2024)
por: Chang, Xuankai, et al.
Publicado: (2024)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
por: Shi, Jiatong, et al.
Publicado: (2024)
por: Shi, Jiatong, et al.
Publicado: (2024)
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
por: Richter, Julius, et al.
Publicado: (2024)
por: Richter, Julius, et al.
Publicado: (2024)
An Empirical Recipe for Universal Phone Recognition
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
por: Tian, Jinchuan, et al.
Publicado: (2024)
por: Tian, Jinchuan, et al.
Publicado: (2024)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
por: Chen, Peikun, et al.
Publicado: (2024)
por: Chen, Peikun, et al.
Publicado: (2024)
HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial Networks
por: Dmitrieva, Ekaterina, et al.
Publicado: (2025)
por: Dmitrieva, Ekaterina, et al.
Publicado: (2025)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
por: Choi, Kwanghee, et al.
Publicado: (2026)
por: Choi, Kwanghee, et al.
Publicado: (2026)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
por: Choi, Kwanghee, et al.
Publicado: (2026)
por: Choi, Kwanghee, et al.
Publicado: (2026)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
por: Maiti, Soumi, et al.
Publicado: (2023)
por: Maiti, Soumi, et al.
Publicado: (2023)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
por: Lee, Philip H., et al.
Publicado: (2024)
por: Lee, Philip H., et al.
Publicado: (2024)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
por: Sun, Haitong, et al.
Publicado: (2026)
por: Sun, Haitong, et al.
Publicado: (2026)
Drax: Speech Recognition with Discrete Flow Matching
por: Navon, Aviv, et al.
Publicado: (2025)
por: Navon, Aviv, et al.
Publicado: (2025)
SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
por: Wang, Pu, et al.
Publicado: (2026)
por: Wang, Pu, et al.
Publicado: (2026)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
por: Bharadwaj, Shikhar, et al.
Publicado: (2025)
por: Bharadwaj, Shikhar, et al.
Publicado: (2025)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
por: Shon, Suwon, et al.
Publicado: (2024)
por: Shon, Suwon, et al.
Publicado: (2024)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
por: Lanzendörfer, Luca A., et al.
Publicado: (2025)
por: Lanzendörfer, Luca A., et al.
Publicado: (2025)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
por: Zeineldeen, Mohammad, et al.
Publicado: (2023)
por: Zeineldeen, Mohammad, et al.
Publicado: (2023)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
por: Ulgen, Ismail Rasim, et al.
Publicado: (2025)
por: Ulgen, Ismail Rasim, et al.
Publicado: (2025)
Speech Watermarking with Discrete Intermediate Representations
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
por: Li, Chin-Jou, et al.
Publicado: (2025)
por: Li, Chin-Jou, et al.
Publicado: (2025)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
por: Huang, Chien-yu, et al.
Publicado: (2023)
por: Huang, Chien-yu, et al.
Publicado: (2023)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
por: Tao, Dehua, et al.
Publicado: (2024)
por: Tao, Dehua, et al.
Publicado: (2024)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
por: Duret, Jarod, et al.
Publicado: (2024)
por: Duret, Jarod, et al.
Publicado: (2024)
Uni-VERSA: Versatile Speech Assessment with a Unified Network
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
por: Sun, Esther, et al.
Publicado: (2026)
por: Sun, Esther, et al.
Publicado: (2026)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
por: Bai, Richard He, et al.
Publicado: (2025)
por: Bai, Richard He, et al.
Publicado: (2025)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
por: Udupa, Sathvik, et al.
Publicado: (2025)
por: Udupa, Sathvik, et al.
Publicado: (2025)
RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
por: Kim, Seungmin, et al.
Publicado: (2025)
por: Kim, Seungmin, et al.
Publicado: (2025)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
por: Saeki, Takaaki, et al.
Publicado: (2024)
por: Saeki, Takaaki, et al.
Publicado: (2024)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
por: Li, Chenda, et al.
Publicado: (2024)
por: Li, Chenda, et al.
Publicado: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
por: Zhang, Wangyou, et al.
Publicado: (2024)
por: Zhang, Wangyou, et al.
Publicado: (2024)
Speech to Speech Synthesis for Voice Impersonation
por: Johnson, Bjorn, et al.
Publicado: (2026)
por: Johnson, Bjorn, et al.
Publicado: (2026)
Ejemplares similares
-
Discrete Speech Unit Extraction via Independent Component Analysis
por: Nakamura, Tomohiko, et al.
Publicado: (2025) -
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
por: Someki, Masao, et al.
Publicado: (2023) -
ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
por: Someki, Masao, et al.
Publicado: (2024) -
Multi-blank Transducers for Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2022) -
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
por: Chang, Xuankai, et al.
Publicado: (2024)