Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Martín-Cortinas, Álvaro, Sáez-Trigueros, Daniel, Vallés-Pérez, Iván, Tura-Vecino, Biel, Biliński, Piotr, Lajszczak, Mateusz, Beringer, Grzegorz, Barra-Chicote, Roberto, Lorenzo-Trueba, Jaime |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating self-supervised features for expressive, multilingual voice conversion
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2025)
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2025)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
von: Guo, Xin, et al.
Veröffentlicht: (2026)
von: Guo, Xin, et al.
Veröffentlicht: (2026)
Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
von: Hanilçi, Cemal, et al.
Veröffentlicht: (2026)
von: Hanilçi, Cemal, et al.
Veröffentlicht: (2026)
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
von: Łajszczak, Mateusz, et al.
Veröffentlicht: (2024)
von: Łajszczak, Mateusz, et al.
Veröffentlicht: (2024)
Information Retrieval for ZeroSpeech 2021: The Submission by University of Wroclaw
von: Chorowski, Jan, et al.
Veröffentlicht: (2021)
von: Chorowski, Jan, et al.
Veröffentlicht: (2021)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
Weakly Supervised Phonological Features for Pathological Speech Analysis
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2025)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2025)
Rethinking Mamba in Speech Processing by Self-Supervised Models
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
Semi-Supervised Contrastive Learning of Musical Representations
von: Guinot, Julien, et al.
Veröffentlicht: (2024)
von: Guinot, Julien, et al.
Veröffentlicht: (2024)
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech
von: Czyżnikiewicz, Mateusz, et al.
Veröffentlicht: (2024)
von: Czyżnikiewicz, Mateusz, et al.
Veröffentlicht: (2024)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
von: Jia, Kaimeng, et al.
Veröffentlicht: (2025)
von: Jia, Kaimeng, et al.
Veröffentlicht: (2025)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
Multi-Distillation from Speech and Music Representation Models
von: Wei, Jui-Chiang, et al.
Veröffentlicht: (2025)
von: Wei, Jui-Chiang, et al.
Veröffentlicht: (2025)
Leveraging Content and Acoustic Representations for Speech Emotion Recognition
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Enhanced Reverberation as Supervision for Unsupervised Speech Separation
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
von: Girish, et al.
Veröffentlicht: (2026)
von: Girish, et al.
Veröffentlicht: (2026)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation
von: Baroudi, Séverin, et al.
Veröffentlicht: (2025)
von: Baroudi, Séverin, et al.
Veröffentlicht: (2025)
TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Learning Time-Graph Frequency Representation for Monaural Speech Enhancement
von: Wang, Tingting, et al.
Veröffentlicht: (2025)
von: Wang, Tingting, et al.
Veröffentlicht: (2025)
ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals
von: E, Ameenudeen P, et al.
Veröffentlicht: (2026)
von: E, Ameenudeen P, et al.
Veröffentlicht: (2026)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Investigating self-supervised features for expressive, multilingual voice conversion
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2025) -
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023) -
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025) -
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025) -
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
von: Guo, Xin, et al.
Veröffentlicht: (2026)