Discrete Speech Unit Extraction via Independent Component Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Nakamura, Tomohiko, Choi, Kwanghee, Hojo, Keigo, Bando, Yoshiaki, Fukayama, Satoru, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On-device Streaming Discrete Speech Units
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
by: Bando, Yoshiaki, et al.
Published: (2024)
by: Bando, Yoshiaki, et al.
Published: (2024)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
by: Bharadwaj, Shikhar, et al.
Published: (2025)
by: Bharadwaj, Shikhar, et al.
Published: (2025)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
by: Bharadwaj, Shikhar, et al.
Published: (2026)
by: Bharadwaj, Shikhar, et al.
Published: (2026)
FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs
by: Suda, Hitoshi, et al.
Published: (2024)
by: Suda, Hitoshi, et al.
Published: (2024)
Self-Supervised Speech Representations are More Phonetic than Semantic
by: Choi, Kwanghee, et al.
Published: (2024)
by: Choi, Kwanghee, et al.
Published: (2024)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
by: Suda, Hitoshi, et al.
Published: (2025)
by: Suda, Hitoshi, et al.
Published: (2025)
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
by: Suda, Hitoshi, et al.
Published: (2025)
by: Suda, Hitoshi, et al.
Published: (2025)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
by: Nakamura, Tomohiko, et al.
Published: (2022)
by: Nakamura, Tomohiko, et al.
Published: (2022)
Real-time Speech Extraction Using Spatially Regularized Independent Low-rank Matrix Analysis and Rank-constrained Spatial Covariance Matrix Estimation
by: Ishikawa, Yuto, et al.
Published: (2024)
by: Ishikawa, Yuto, et al.
Published: (2024)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
by: Chang, Xuankai, et al.
Published: (2024)
by: Chang, Xuankai, et al.
Published: (2024)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
by: Tian, Jinchuan, et al.
Published: (2024)
by: Tian, Jinchuan, et al.
Published: (2024)
Head-Related Transfer Function Individualization Using Anthropometric Features and Spatially Independent Latent Representation
by: Niu, Ryan, et al.
Published: (2025)
by: Niu, Ryan, et al.
Published: (2025)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
by: Wang, Yongqi, et al.
Published: (2023)
by: Wang, Yongqi, et al.
Published: (2023)
Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
by: Poncelet, Jakob, et al.
Published: (2024)
by: Poncelet, Jakob, et al.
Published: (2024)
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
by: Koga, Naoki, et al.
Published: (2024)
by: Koga, Naoki, et al.
Published: (2024)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
by: Shon, Suwon, et al.
Published: (2024)
by: Shon, Suwon, et al.
Published: (2024)
Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides
by: Imamura, Kanami, et al.
Published: (2023)
by: Imamura, Kanami, et al.
Published: (2023)
ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
by: Someki, Masao, et al.
Published: (2024)
by: Someki, Masao, et al.
Published: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
by: Wang, Shih-heng, et al.
Published: (2024)
by: Wang, Shih-heng, et al.
Published: (2024)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
by: Shi, Jiatong, et al.
Published: (2024)
by: Shi, Jiatong, et al.
Published: (2024)
Local Equivariance Error-Based Metrics for Evaluating Sampling-Frequency-Independent Property of Neural Network
by: Imamura, Kanami, et al.
Published: (2025)
by: Imamura, Kanami, et al.
Published: (2025)
Infrastructure-less Localization from Indoor Environmental Sounds Based on Spectral Decomposition and Spatial Likelihood Model
by: Ogiso, Satoki, et al.
Published: (2024)
by: Ogiso, Satoki, et al.
Published: (2024)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
by: Choi, Yerin, et al.
Published: (2024)
by: Choi, Yerin, et al.
Published: (2024)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
by: Maekaku, Takashi, et al.
Published: (2025)
by: Maekaku, Takashi, et al.
Published: (2025)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
by: Tao, Dehua, et al.
Published: (2024)
by: Tao, Dehua, et al.
Published: (2024)
Uni-VERSA: Versatile Speech Assessment with a Unified Network
by: Shi, Jiatong, et al.
Published: (2025)
by: Shi, Jiatong, et al.
Published: (2025)
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
by: Nishikawa, Go, et al.
Published: (2025)
by: Nishikawa, Go, et al.
Published: (2025)
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
by: Wang, Shih-Heng, et al.
Published: (2026)
by: Wang, Shih-Heng, et al.
Published: (2026)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
by: Saeki, Takaaki, et al.
Published: (2024)
by: Saeki, Takaaki, et al.
Published: (2024)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
by: Li, Chenda, et al.
Published: (2024)
by: Li, Chenda, et al.
Published: (2024)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
by: Chen, Peikun, et al.
Published: (2024)
by: Chen, Peikun, et al.
Published: (2024)
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
by: Shi, Jiatong, et al.
Published: (2025)
by: Shi, Jiatong, et al.
Published: (2025)
Improving Design of Input Condition Invariant Speech Enhancement
by: Zhang, Wangyou, et al.
Published: (2024)
by: Zhang, Wangyou, et al.
Published: (2024)
Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
by: Wada, Aogu, et al.
Published: (2025)
by: Wada, Aogu, et al.
Published: (2025)
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
by: Fujita, Yoto, et al.
Published: (2024)
by: Fujita, Yoto, et al.
Published: (2024)
Differentiable K-means for Fully-optimized Discrete Token-based ASR
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
Similar Items
-
On-device Streaming Discrete Speech Units
by: Choi, Kwanghee, et al.
Published: (2025) -
Neural Blind Source Separation and Diarization for Distant Speech Recognition
by: Bando, Yoshiaki, et al.
Published: (2024) -
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
by: Bharadwaj, Shikhar, et al.
Published: (2025) -
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
by: Bharadwaj, Shikhar, et al.
Published: (2026) -
FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs
by: Suda, Hitoshi, et al.
Published: (2024)