Guardado en:
| Autores principales: | Naderi, Maryam, Hermann, Enno, Nanchen, Alexandre, Hovsepyan, Sevada, -Doss, Mathew Magimai. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2407.21414 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
por: Sarkar, Eklavya, et al.
Publicado: (2025)
por: Sarkar, Eklavya, et al.
Publicado: (2025)
kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech
por: Hajal, Karl El, et al.
Publicado: (2024)
por: Hajal, Karl El, et al.
Publicado: (2024)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
por: Sarkar, Eklavya, et al.
Publicado: (2024)
por: Sarkar, Eklavya, et al.
Publicado: (2024)
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
por: Elbanna, Gasser, et al.
Publicado: (2024)
por: Elbanna, Gasser, et al.
Publicado: (2024)
Toward using Speech to Sense Student Emotion in Remote Learning Environments
por: Vyas, Sargam, et al.
Publicado: (2026)
por: Vyas, Sargam, et al.
Publicado: (2026)
On feature representations for marmoset vocal communication analysis
por: Sarkar, Eklavya, et al.
Publicado: (2025)
por: Sarkar, Eklavya, et al.
Publicado: (2025)
Feature Representations for Automatic Meerkat Vocalization Classification
por: Mahmoud, Imen Ben, et al.
Publicado: (2024)
por: Mahmoud, Imen Ben, et al.
Publicado: (2024)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
por: Zhang, Alice, et al.
Publicado: (2025)
por: Zhang, Alice, et al.
Publicado: (2025)
Building English ASR model with regional language support
por: Agrawal, Purvi, et al.
Publicado: (2025)
por: Agrawal, Purvi, et al.
Publicado: (2025)
Extending Whisper with prompt tuning to target-speaker ASR
por: Ma, Hao, et al.
Publicado: (2023)
por: Ma, Hao, et al.
Publicado: (2023)
Supplementary Information for: "Speech power spectra: a window into neural oscillations in Parkinson's disease"
por: Hovsepyan, Sevada, et al.
Publicado: (2025)
por: Hovsepyan, Sevada, et al.
Publicado: (2025)
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
por: Dowerah, Sandipana, et al.
Publicado: (2025)
por: Dowerah, Sandipana, et al.
Publicado: (2025)
Towards scalable efficient on-device ASR with transfer learning
por: Pandey, Laxmi, et al.
Publicado: (2024)
por: Pandey, Laxmi, et al.
Publicado: (2024)
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
por: Kashiwagi, Yosuke, et al.
Publicado: (2025)
por: Kashiwagi, Yosuke, et al.
Publicado: (2025)
Transferable speech-to-text large language model alignment module
por: Wu, Boyong, et al.
Publicado: (2024)
por: Wu, Boyong, et al.
Publicado: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
por: Xie, Yuan, et al.
Publicado: (2026)
por: Xie, Yuan, et al.
Publicado: (2026)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
por: Hernandez, Abner, et al.
Publicado: (2026)
por: Hernandez, Abner, et al.
Publicado: (2026)
Robustness assessment of large audio language models in multiple-choice evaluation
por: López, Fernando, et al.
Publicado: (2025)
por: López, Fernando, et al.
Publicado: (2025)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
por: Li, Chia-Yu, et al.
Publicado: (2024)
por: Li, Chia-Yu, et al.
Publicado: (2024)
Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
por: Patil, Aditya, et al.
Publicado: (2024)
por: Patil, Aditya, et al.
Publicado: (2024)
PromptASR for contextualized ASR with controllable style
por: Yang, Xiaoyu, et al.
Publicado: (2023)
por: Yang, Xiaoyu, et al.
Publicado: (2023)
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
por: Everson, Kevin, et al.
Publicado: (2024)
por: Everson, Kevin, et al.
Publicado: (2024)
The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
por: Chen, William, et al.
Publicado: (2025)
por: Chen, William, et al.
Publicado: (2025)
Unified Learnable 2D Convolutional Feature Extraction for ASR
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
ASR Error Correction using Large Language Models
por: Ma, Rao, et al.
Publicado: (2024)
por: Ma, Rao, et al.
Publicado: (2024)
Word-wise intonation model for cross-language TTS systems
por: A., Tomilov A., et al.
Publicado: (2024)
por: A., Tomilov A., et al.
Publicado: (2024)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
por: Kwok, Chin Yuen, et al.
Publicado: (2024)
por: Kwok, Chin Yuen, et al.
Publicado: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
por: Okocha, Chibuzor, et al.
Publicado: (2025)
por: Okocha, Chibuzor, et al.
Publicado: (2025)
Improving ASR Contextual Biasing with Guided Attention
por: Tang, Jiyang, et al.
Publicado: (2024)
por: Tang, Jiyang, et al.
Publicado: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
por: Gündüz, Ahmet, et al.
Publicado: (2024)
por: Gündüz, Ahmet, et al.
Publicado: (2024)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
por: Kumar, Shashi, et al.
Publicado: (2024)
por: Kumar, Shashi, et al.
Publicado: (2024)
Towards continually learning new languages
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
Robust ASR Error Correction with Conservative Data Filtering
por: Udagawa, Takuma, et al.
Publicado: (2024)
por: Udagawa, Takuma, et al.
Publicado: (2024)
Retrieval Augmented Generation based context discovery for ASR
por: Siskos, Dimitrios, et al.
Publicado: (2025)
por: Siskos, Dimitrios, et al.
Publicado: (2025)
Optimizing Byte-level Representation for End-to-end ASR
por: Hsiao, Roger, et al.
Publicado: (2024)
por: Hsiao, Roger, et al.
Publicado: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
Ejemplares similares
-
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
por: Hajal, Karl El, et al.
Publicado: (2025) -
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
por: Hajal, Karl El, et al.
Publicado: (2025) -
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
por: Sarkar, Eklavya, et al.
Publicado: (2025) -
kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech
por: Hajal, Karl El, et al.
Publicado: (2024) -
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
por: Sarkar, Eklavya, et al.
Publicado: (2024)