CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Yen-Ju, Liu, Jing, Thebaud, Thomas, Moro-Velazquez, Laureano, Rastrow, Ariya, Dehak, Najim, Villalba, Jesus |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
von: Joshi, Sonal, et al.
Veröffentlicht: (2021)
von: Joshi, Sonal, et al.
Veröffentlicht: (2021)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
Noise-robust Speech Separation with Fast Generative Correction
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification
von: Joshi, Sonal, et al.
Veröffentlicht: (2024)
von: Joshi, Sonal, et al.
Veröffentlicht: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
von: Frummer, Ari, et al.
Veröffentlicht: (2025)
von: Frummer, Ari, et al.
Veröffentlicht: (2025)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
von: Thebaud, Thomas, et al.
Veröffentlicht: (2025)
von: Thebaud, Thomas, et al.
Veröffentlicht: (2025)
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
von: Lee, Junhyeok, et al.
Veröffentlicht: (2026)
von: Lee, Junhyeok, et al.
Veröffentlicht: (2026)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
Adversarial Attacks and Defenses for Speech Recognition Systems
von: Żelasko, Piotr, et al.
Veröffentlicht: (2021)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2021)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
von: Lee, Junhyeok, et al.
Veröffentlicht: (2025)
von: Lee, Junhyeok, et al.
Veröffentlicht: (2025)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Additive Margin in Contrastive Self-Supervised Frameworks to Learn Discriminative Speaker Representations
von: Lepage, Theo, et al.
Veröffentlicht: (2024)
von: Lepage, Theo, et al.
Veröffentlicht: (2024)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
von: Thebaud, Thomas, et al.
Veröffentlicht: (2026)
von: Thebaud, Thomas, et al.
Veröffentlicht: (2026)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Two-pass Endpoint Detection for Speech Recognition
von: Raju, Anirudh, et al.
Veröffentlicht: (2024)
von: Raju, Anirudh, et al.
Veröffentlicht: (2024)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
Self-Supervised Learning for Speaker Recognition: A study and review
von: Lepage, Theo, et al.
Veröffentlicht: (2026)
von: Lepage, Theo, et al.
Veröffentlicht: (2026)
Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
von: Lepage, Théo, et al.
Veröffentlicht: (2022)
von: Lepage, Théo, et al.
Veröffentlicht: (2022)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
von: Miara, Victor, et al.
Veröffentlicht: (2024)
von: Miara, Victor, et al.
Veröffentlicht: (2024)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
Automatic Proficiency Assessment in L2 English Learners
von: Mohammadi, Armita, et al.
Veröffentlicht: (2025)
von: Mohammadi, Armita, et al.
Veröffentlicht: (2025)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
von: Joshi, Sonal, et al.
Veröffentlicht: (2021) -
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
von: Wang, Helin, et al.
Veröffentlicht: (2025) -
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026) -
Noise-robust Speech Separation with Fast Generative Correction
von: Wang, Helin, et al.
Veröffentlicht: (2024) -
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification
von: Joshi, Sonal, et al.
Veröffentlicht: (2024)