persoDA: Personalized Data Augmentation for Personalized ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Parada, Pablo Peso, Fontalis, Spyros, Jalal, Md Asif, Saravanan, Karthikeyan, Drosou, Anastasios, Ozay, Mete, Lee, Gil Ho, Lee, Jungin, Jung, Seokyeong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025)
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025)
Locality enhanced dynamic biasing and sampling strategies for contextual ASR
von: Jalal, Md Asif, et al.
Veröffentlicht: (2024)
von: Jalal, Md Asif, et al.
Veröffentlicht: (2024)
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
Retrieval Augmented Generation based context discovery for ASR
von: Siskos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Siskos, Dimitrios, et al.
Veröffentlicht: (2025)
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
von: Zhang, Jisi, et al.
Veröffentlicht: (2025)
von: Zhang, Jisi, et al.
Veröffentlicht: (2025)
Exploring compressibility of transformer based text-to-music (TTM) models
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2026)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2026)
A Model for Every User and Budget: Label-Free and Personalized Mixed-Precision Quantization
von: Fish, Edward, et al.
Veröffentlicht: (2023)
von: Fish, Edward, et al.
Veröffentlicht: (2023)
DisAgg: Distributed Aggregators for Efficient Secure Aggregation in Federated Learning
von: Mehmood, Haaris, et al.
Veröffentlicht: (2026)
von: Mehmood, Haaris, et al.
Veröffentlicht: (2026)
Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling
von: Jalal, Md Asif, et al.
Veröffentlicht: (2025)
von: Jalal, Md Asif, et al.
Veröffentlicht: (2025)
Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
Personalized Voice Synthesis through Human-in-the-Loop Coordinate Descent
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2024)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2024)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
von: Wu, Ya-Tse, et al.
Veröffentlicht: (2026)
von: Wu, Ya-Tse, et al.
Veröffentlicht: (2026)
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
von: Jin, Zengrui, et al.
Veröffentlicht: (2022)
von: Jin, Zengrui, et al.
Veröffentlicht: (2022)
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
von: Feng, Fuyuan, et al.
Veröffentlicht: (2026)
von: Feng, Fuyuan, et al.
Veröffentlicht: (2026)
Adversarial speech for voice privacy protection from Personalized Speech generation
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
Causal Analysis of ASR Errors for Children: Quantifying the Impact of Physiological, Cognitive, and Extrinsic Factors
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
ManWav: The First Manchu ASR Model
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR
von: Ogun, Sewade, et al.
Veröffentlicht: (2025)
von: Ogun, Sewade, et al.
Veröffentlicht: (2025)
A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
von: Yen, Hao, et al.
Veröffentlicht: (2025)
von: Yen, Hao, et al.
Veröffentlicht: (2025)
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
von: Lee, Chia-Yu, et al.
Veröffentlicht: (2026)
von: Lee, Chia-Yu, et al.
Veröffentlicht: (2026)
A Parameter-efficient Language Extension Framework for Multilingual ASR
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization
von: Tabib, H. M. Shadman, et al.
Veröffentlicht: (2026)
von: Tabib, H. M. Shadman, et al.
Veröffentlicht: (2026)
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
von: Bae, Jae-Sung, et al.
Veröffentlicht: (2025)
von: Bae, Jae-Sung, et al.
Veröffentlicht: (2025)
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
Causal Structure Discovery for Error Diagnostics of Children's ASR
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
SAM: A Mamba-2 State-Space Audio-Language Model
von: Lee, Taehan, et al.
Veröffentlicht: (2025)
von: Lee, Taehan, et al.
Veröffentlicht: (2025)
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
von: Huang, Shao-Syuan, et al.
Veröffentlicht: (2024)
von: Huang, Shao-Syuan, et al.
Veröffentlicht: (2024)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025) -
Locality enhanced dynamic biasing and sampling strategies for contextual ASR
von: Jalal, Md Asif, et al.
Veröffentlicht: (2024) -
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024) -
Retrieval Augmented Generation based context discovery for ASR
von: Siskos, Dimitrios, et al.
Veröffentlicht: (2025) -
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
von: Zhang, Jisi, et al.
Veröffentlicht: (2025)