Factorized RVQ-GAN For Disentangled Speech Tokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Khurana, Sameer, Klement, Dominik, Laurent, Antoine, Bobos, Dominik, Novosad, Juraj, Gazdik, Peter, Zhang, Ellen, Huang, Zili, Hussein, Amir, Marxer, Ricard, Masuyama, Yoshiki, Aihara, Ryo, Hori, Chiori, Germain, Francois G., Wichern, Gordon, Roux, Jonathan Le |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
by: Masuyama, Yoshiki, et al.
Published: (2024)
by: Masuyama, Yoshiki, et al.
Published: (2024)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
by: Aihara, Ryo, et al.
Published: (2025)
by: Aihara, Ryo, et al.
Published: (2025)
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
by: Masuyama, Yoshiki, et al.
Published: (2026)
by: Masuyama, Yoshiki, et al.
Published: (2026)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
by: Hussein, Amir, et al.
Published: (2025)
by: Hussein, Amir, et al.
Published: (2025)
SUNAC: Source-aware Unified Neural Audio Codec
by: Aihara, Ryo, et al.
Published: (2025)
by: Aihara, Ryo, et al.
Published: (2025)
FasTUSS: Faster Task-Aware Unified Source Separation
by: Paissan, Francesco, et al.
Published: (2025)
by: Paissan, Francesco, et al.
Published: (2025)
Direction-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses
by: Ick, Christopher, et al.
Published: (2025)
by: Ick, Christopher, et al.
Published: (2025)
Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
by: Ick, Christopher, et al.
Published: (2025)
by: Ick, Christopher, et al.
Published: (2025)
Physics-Informed Direction-Aware Neural Acoustic Fields
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
by: Koo, Junghyun, et al.
Published: (2024)
by: Koo, Junghyun, et al.
Published: (2024)
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
by: Saijo, Kohei, et al.
Published: (2024)
by: Saijo, Kohei, et al.
Published: (2024)
Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
by: Hori, Chiori, et al.
Published: (2025)
by: Hori, Chiori, et al.
Published: (2025)
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
by: Richter, Julius, et al.
Published: (2026)
by: Richter, Julius, et al.
Published: (2026)
Aligning Multimodal Representations through an Information Bottleneck
by: Almudévar, Antonio, et al.
Published: (2025)
by: Almudévar, Antonio, et al.
Published: (2025)
Sound Event Bounding Boxes
by: Ebbers, Janek, et al.
Published: (2024)
by: Ebbers, Janek, et al.
Published: (2024)
Why does music source separation benefit from cacophony?
by: Jeon, Chang-Bin, et al.
Published: (2024)
by: Jeon, Chang-Bin, et al.
Published: (2024)
Enhanced Reverberation as Supervision for Unsupervised Speech Separation
by: Saijo, Kohei, et al.
Published: (2024)
by: Saijo, Kohei, et al.
Published: (2024)
Task-Aware Unified Source Separation
by: Saijo, Kohei, et al.
Published: (2024)
by: Saijo, Kohei, et al.
Published: (2024)
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
by: Saijo, Kohei, et al.
Published: (2024)
by: Saijo, Kohei, et al.
Published: (2024)
Speech foundation models on intelligibility prediction for hearing-impaired listeners
by: Cuervo, Santiago, et al.
Published: (2024)
by: Cuervo, Santiago, et al.
Published: (2024)
Scaling Properties of Speech Language Models
by: Cuervo, Santiago, et al.
Published: (2024)
by: Cuervo, Santiago, et al.
Published: (2024)
Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs
by: Cuervo, Santiago, et al.
Published: (2025)
by: Cuervo, Santiago, et al.
Published: (2025)
Local Density-Based Anomaly Score Normalization for Domain Generalization
by: Wilkinghoff, Kevin, et al.
Published: (2025)
by: Wilkinghoff, Kevin, et al.
Published: (2025)
Transfer Learning from Whisper for Microscopic Intelligibility Prediction
by: Best, Paul, et al.
Published: (2024)
by: Best, Paul, et al.
Published: (2024)
Robust Training of Vector Quantized Bottleneck Models
by: Łańcucki, Adrian, et al.
Published: (2020)
by: Łańcucki, Adrian, et al.
Published: (2020)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
by: Wu, Shih-Lun, et al.
Published: (2023)
by: Wu, Shih-Lun, et al.
Published: (2023)
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
Exploring the Capability of Mamba in Speech Applications
by: Miyazaki, Koichi, et al.
Published: (2024)
by: Miyazaki, Koichi, et al.
Published: (2024)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
by: Masuyama, Yoshiki, et al.
Published: (2024)
by: Masuyama, Yoshiki, et al.
Published: (2024)
Depth Jitter: Seeing through the Depth
by: Rahman, Md Sazidur, et al.
Published: (2025)
by: Rahman, Md Sazidur, et al.
Published: (2025)
Mind the Gap: Detecting Cluster Exits for Robust Local Density-Based Score Normalization in Anomalous Sound Detection
by: Wilkinghoff, Kevin, et al.
Published: (2026)
by: Wilkinghoff, Kevin, et al.
Published: (2026)
On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation
by: Baroudi, Séverin, et al.
Published: (2025)
by: Baroudi, Séverin, et al.
Published: (2025)
Deep Learning Classification With Noisy Labels
by: Sanchez, Guillaume, et al.
Published: (2020)
by: Sanchez, Guillaume, et al.
Published: (2020)
Applying machine learning to primate bioacoustics: Review and perspectives
by: Jules Cauzinille, et al.
Published: (2024)
by: Jules Cauzinille, et al.
Published: (2024)
Dictionnaire historique de l’adjectif-adverbe, Volume 1
by: Hummel, Martin, et al.
Published: (2021)
by: Hummel, Martin, et al.
Published: (2021)
Dictionnaire historique de l’adjectif-adverbe, Volume 2
by: Hummel, Martin, et al.
Published: (2021)
by: Hummel, Martin, et al.
Published: (2021)
Dictionnaire historique de l’adjectif-adverbe
by: Hummel, Martin, et al.
Published: (2022)
by: Hummel, Martin, et al.
Published: (2022)
Unsupervised Speech Enhancement using Data-defined Priors
by: Klement, Dominik, et al.
Published: (2025)
by: Klement, Dominik, et al.
Published: (2025)
Similar Items
-
NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
by: Masuyama, Yoshiki, et al.
Published: (2024) -
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
by: Aihara, Ryo, et al.
Published: (2025) -
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
by: Masuyama, Yoshiki, et al.
Published: (2026) -
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
by: Hussein, Amir, et al.
Published: (2025) -
SUNAC: Source-aware Unified Neural Audio Codec
by: Aihara, Ryo, et al.
Published: (2025)