ZIPA: A family of efficient models for multilingual phone recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Jian, Samir, Farhan, Chodroff, Eleanor, Mortensen, David R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
von: Zhu, Jian, et al.
Veröffentlicht: (2023)
von: Zhu, Jian, et al.
Veröffentlicht: (2023)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
A two-stage transliteration approach to improve performance of a multilingual ASR
von: Kumar, Rohit
Veröffentlicht: (2024)
von: Kumar, Rohit
Veröffentlicht: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
von: Mei, Yuxiang, et al.
Veröffentlicht: (2026)
von: Mei, Yuxiang, et al.
Veröffentlicht: (2026)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversation emotion recognition
von: Liu, Ying, et al.
Veröffentlicht: (2026)
von: Liu, Ying, et al.
Veröffentlicht: (2026)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
fastabx: A library for efficient computation of ABX discriminability
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
von: You, Jian, et al.
Veröffentlicht: (2024)
von: You, Jian, et al.
Veröffentlicht: (2024)
Adjust-free adversarial example generation in speech recognition using evolutionary multi-objective optimization under black-box condition
von: Ishida, Shoma, et al.
Veröffentlicht: (2020)
von: Ishida, Shoma, et al.
Veröffentlicht: (2020)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
A multilingual training strategy for low resource Text to Speech
von: Amalas, Asma, et al.
Veröffentlicht: (2024)
von: Amalas, Asma, et al.
Veröffentlicht: (2024)
Semantic enrichment towards efficient speech representations
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
TidyVoice 2026 Challenge Evaluation Plan
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
AUDDT: Audio Unified Deepfake Detection Benchmark Toolkit
von: Zhu, Yi, et al.
Veröffentlicht: (2025)
von: Zhu, Yi, et al.
Veröffentlicht: (2025)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
von: Dat, Phuong Tuan, et al.
Veröffentlicht: (2025)
von: Dat, Phuong Tuan, et al.
Veröffentlicht: (2025)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
Alethia: A Foundational Encoder for Voice Deepfakes
von: Zhu, Yi, et al.
Veröffentlicht: (2026)
von: Zhu, Yi, et al.
Veröffentlicht: (2026)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024) -
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
von: Zhu, Jian, et al.
Veröffentlicht: (2023) -
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026) -
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026) -
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)