POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Chin-Jou, Chang, Kalvin, Bharadwaj, Shikhar, Yeo, Eunjung, Choi, Kwanghee, Zhu, Jian, Mortensen, David, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
An Empirical Recipe for Universal Phone Recognition
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Representations are More Phonetic than Semantic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
On-device Streaming Discrete Speech Units
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
von: Yeo, Eunjung
Veröffentlicht: (2024)
von: Yeo, Eunjung
Veröffentlicht: (2024)
Context-Driven Dynamic Pruning for Large Speech Foundation Models
von: Someki, Masao, et al.
Veröffentlicht: (2025)
von: Someki, Masao, et al.
Veröffentlicht: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
Visual Speech Recognition for Languages with Limited Labeled Data using Automatic Labels from Whisper
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
Discrete Speech Unit Extraction via Independent Component Analysis
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2025)
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2025)
OpusLM: A Family of Open Unified Speech Language Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
von: Yan, Brian, et al.
Veröffentlicht: (2025)
von: Yan, Brian, et al.
Veröffentlicht: (2025)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
von: Li, Zirui, et al.
Veröffentlicht: (2025)
von: Li, Zirui, et al.
Veröffentlicht: (2025)
Speech Codec Probing from Semantic and Phonetic Perspectives
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
STAB: Speech Tokenizer Assessment Benchmark
von: Vashishth, Shikhar, et al.
Veröffentlicht: (2024)
von: Vashishth, Shikhar, et al.
Veröffentlicht: (2024)
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Uni-VERSA: Versatile Speech Assessment with a Unified Network
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2026)
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2026)
Prompting Whisper for Joint Speech Transcription and Diarization
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
Probing Whisper for Dysarthric Speech in Detection and Assessment
von: Yue, Zhengjun, et al.
Veröffentlicht: (2025)
von: Yue, Zhengjun, et al.
Veröffentlicht: (2025)
AS-Speech: Adaptive Style For Speech Synthesis
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025) -
An Empirical Recipe for Universal Phone Recognition
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026) -
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026) -
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025) -
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)