Exploring Dynamic Parameters for Vietnamese Gender-Independent ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Leang, Sotheara, Castelli, Éric, Vaufreydaz, Dominique, Sam, Sethserey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization
von: Leang, Sotheara, et al.
Veröffentlicht: (2024)
von: Leang, Sotheara, et al.
Veröffentlicht: (2024)
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
An ASR-Based Tutor for Learning to Read: How to Optimize Feedback to First Graders
von: Bai, Yu, et al.
Veröffentlicht: (2023)
von: Bai, Yu, et al.
Veröffentlicht: (2023)
Hybrid Deep Learning and Signal Processing for Arabic Dialect Recognition in Low-Resource Settings
von: Al-Shwayyat, Ghazal, et al.
Veröffentlicht: (2025)
von: Al-Shwayyat, Ghazal, et al.
Veröffentlicht: (2025)
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics
von: Molenaar, Bo, et al.
Veröffentlicht: (2023)
von: Molenaar, Bo, et al.
Veröffentlicht: (2023)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
von: Wills, Simone, et al.
Veröffentlicht: (2023)
von: Wills, Simone, et al.
Veröffentlicht: (2023)
Exploring SSL Discrete Tokens for Multilingual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Interpolation Filter Design for Sample Rate Independent Audio Effect RNNs
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
Sample Rate Independent Recurrent Neural Networks for Audio Effects Processing
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
Online Similarity-and-Independence-Aware Beamformer for Low-latency Target Sound Extraction
von: Hiroe, Atsuo
Veröffentlicht: (2023)
von: Hiroe, Atsuo
Veröffentlicht: (2023)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
Exploring Gender Disparities in Automatic Speech Recognition Technology
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
Promptformer: Prompted Conformer Transducer for ASR
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
Qwen3-ASR Technical Report
von: Shi, Xian, et al.
Veröffentlicht: (2026)
von: Shi, Xian, et al.
Veröffentlicht: (2026)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Dynamic Prediction of Full-Ocean Depth SSP by Hierarchical LSTM: An Experimental Result
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
ManWav: The First Manchu ASR Model
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024)
von: Fang, Ying, et al.
Veröffentlicht: (2024)
Scalable Offline ASR for Command-Style Dictation in Courtrooms
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
Causal Structure Discovery for Error Diagnostics of Children's ASR
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
Performant ASR Models for Medical Entities in Accented Speech
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
Reverb: Open-Source ASR and Diarization from Rev
von: Bhandari, Nishchal, et al.
Veröffentlicht: (2024)
von: Bhandari, Nishchal, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization
von: Leang, Sotheara, et al.
Veröffentlicht: (2024) -
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023) -
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025) -
An ASR-Based Tutor for Learning to Read: How to Optimize Feedback to First Graders
von: Bai, Yu, et al.
Veröffentlicht: (2023) -
Hybrid Deep Learning and Signal Processing for Arabic Dialect Recognition in Low-Resource Settings
von: Al-Shwayyat, Ghazal, et al.
Veröffentlicht: (2025)