VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space
Fuente:
arXiv
Saved in:
| Main Authors: | Rodriguez, Armani, Kokalj-Filipovic, Silvija |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep-Learned Compression for Radio-Frequency Signal Classification
by: Rodriguez, Armani, et al.
Published: (2024)
by: Rodriguez, Armani, et al.
Published: (2024)
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
by: Li, Jiatong, et al.
Published: (2025)
by: Li, Jiatong, et al.
Published: (2025)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
by: Yang, Yudong, et al.
Published: (2024)
by: Yang, Yudong, et al.
Published: (2024)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
by: Jiang, Ziyue, et al.
Published: (2025)
by: Jiang, Ziyue, et al.
Published: (2025)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
by: Niu, Zhikang, et al.
Published: (2025)
by: Niu, Zhikang, et al.
Published: (2025)
Latent-Domain Predictive Neural Speech Coding
by: Jiang, Xue, et al.
Published: (2022)
by: Jiang, Xue, et al.
Published: (2022)
Latent Speech-Text Transformer
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
by: Lovelace, Justin, et al.
Published: (2025)
by: Lovelace, Justin, et al.
Published: (2025)
Investigating the Effects of Diffusion-based Conditional Generative Speech Models Used for Speech Enhancement on Dysarthric Speech
by: Reszka, Joanna, et al.
Published: (2024)
by: Reszka, Joanna, et al.
Published: (2024)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
by: Varshavsky-Hassid, Miri, et al.
Published: (2024)
by: Varshavsky-Hassid, Miri, et al.
Published: (2024)
Speech Enhancement and Dereverberation with Diffusion-based Generative Models
by: Richter, Julius, et al.
Published: (2022)
by: Richter, Julius, et al.
Published: (2022)
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis
by: Lin, Weiwei, et al.
Published: (2024)
by: Lin, Weiwei, et al.
Published: (2024)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
by: Luong, Diep, et al.
Published: (2025)
by: Luong, Diep, et al.
Published: (2025)
Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
by: Gerazov, Branislav, et al.
Published: (2025)
by: Gerazov, Branislav, et al.
Published: (2025)
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
by: Richter, Julius, et al.
Published: (2026)
by: Richter, Julius, et al.
Published: (2026)
MaskCycleGAN-based Whisper to Normal Speech Conversion
by: Gupta, K. Rohith, et al.
Published: (2024)
by: Gupta, K. Rohith, et al.
Published: (2024)
ReFormer: Generating Radio Fakes for Data Augmentation
by: Kaasaragadda, Yagna, et al.
Published: (2024)
by: Kaasaragadda, Yagna, et al.
Published: (2024)
Investigating the Design Space of Diffusion Models for Speech Enhancement
by: Gonzalez, Philippe, et al.
Published: (2023)
by: Gonzalez, Philippe, et al.
Published: (2023)
Voice Disorder Analysis: a Transformer-based Approach
by: Koudounas, Alkis, et al.
Published: (2024)
by: Koudounas, Alkis, et al.
Published: (2024)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
by: Lee, Hyeonseung, et al.
Published: (2024)
by: Lee, Hyeonseung, et al.
Published: (2024)
EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
by: Cerovaz, Luca, et al.
Published: (2026)
by: Cerovaz, Luca, et al.
Published: (2026)
Diffusion Buffer: Online Diffusion-based Speech Enhancement with Sub-Second Latency
by: Lay, Bunlong, et al.
Published: (2025)
by: Lay, Bunlong, et al.
Published: (2025)
I-DCCRN-VAE: An Improved Deep Representation Learning Framework for Complex VAE-based Single-channel Speech Enhancement
by: Li, Jiatong, et al.
Published: (2025)
by: Li, Jiatong, et al.
Published: (2025)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
by: Narain, Jaya, et al.
Published: (2025)
by: Narain, Jaya, et al.
Published: (2025)
Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
Quartered Spectral Envelope and 1D-CNN-based Classification of Normally Phonated and Whispered Speech
by: Joysingh, S. Johanan, et al.
Published: (2024)
by: Joysingh, S. Johanan, et al.
Published: (2024)
Multi-Source Music Generation with Latent Diffusion
by: Xu, Zhongweiyang, et al.
Published: (2024)
by: Xu, Zhongweiyang, et al.
Published: (2024)
Bass Accompaniment Generation via Latent Diffusion
by: Pasini, Marco, et al.
Published: (2024)
by: Pasini, Marco, et al.
Published: (2024)
SECP: A Speech Enhancement-Based Curation Pipeline For Scalable Acquisition Of Clean Speech
by: Sabra, Adam, et al.
Published: (2024)
by: Sabra, Adam, et al.
Published: (2024)
Diffusion Buffer for Online Generative Speech Enhancement
by: Lay, Bunlong, et al.
Published: (2025)
by: Lay, Bunlong, et al.
Published: (2025)
An Analysis of the Variance of Diffusion-based Speech Enhancement
by: Lay, Bunlong, et al.
Published: (2024)
by: Lay, Bunlong, et al.
Published: (2024)
Towards Audio Codec-based Speech Separation
by: Yip, Jia Qi, et al.
Published: (2024)
by: Yip, Jia Qi, et al.
Published: (2024)
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
by: Cao, Fengyuan, et al.
Published: (2026)
by: Cao, Fengyuan, et al.
Published: (2026)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
by: Yang, Jinhyeok, et al.
Published: (2026)
by: Yang, Jinhyeok, et al.
Published: (2026)
A Context-Based Numerical Format Prediction for a Text-To-Speech System
by: Darwesh, Yaser, et al.
Published: (2024)
by: Darwesh, Yaser, et al.
Published: (2024)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
by: Lou, Haowei, et al.
Published: (2024)
by: Lou, Haowei, et al.
Published: (2024)
Discrete-Time Diffusion-Like Models for Speech Synthesis
by: Tan, Xiaozhou, et al.
Published: (2025)
by: Tan, Xiaozhou, et al.
Published: (2025)
Speech Understanding on Tiny Devices with A Learning Cache
by: Benazir, Afsara, et al.
Published: (2023)
by: Benazir, Afsara, et al.
Published: (2023)
Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
by: Yanuka, Moran, et al.
Published: (2025)
by: Yanuka, Moran, et al.
Published: (2025)
Similar Items
-
Deep-Learned Compression for Radio-Frequency Signal Classification
by: Rodriguez, Armani, et al.
Published: (2024) -
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
by: Li, Jiatong, et al.
Published: (2025) -
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
by: Melechovsky, Jan, et al.
Published: (2024) -
Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
by: Yang, Yudong, et al.
Published: (2024) -
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
by: Jiang, Ziyue, et al.
Published: (2025)