PRiSM: Benchmarking Phone Realization in Speech Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bharadwaj, Shikhar, Li, Chin-Jou, Kim, Yoonjae, Choi, Kwanghee, Yeo, Eunjung, Shim, Ryan Soh-Eun, Zhou, Hanyu, Boldt, Brendon, Jacome, Karen Rosero, Chang, Kalvin, Agrawal, Darsh, Xu, Keer, Yang, Chao-Han Huck, Zhu, Jian, Watanabe, Shinji, Mortensen, David R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Recipe for Universal Phone Recognition
by: Bharadwaj, Shikhar, et al.
Published: (2026)
by: Bharadwaj, Shikhar, et al.
Published: (2026)
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
by: Li, Chin-Jou, et al.
Published: (2025)
by: Li, Chin-Jou, et al.
Published: (2025)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
XferBench: a Data-Driven Benchmark for Emergent Language
by: Boldt, Brendon, et al.
Published: (2024)
by: Boldt, Brendon, et al.
Published: (2024)
Searching for the Most Human-like Emergent Language
by: Boldt, Brendon, et al.
Published: (2025)
by: Boldt, Brendon, et al.
Published: (2025)
A Review of the Applications of Deep Learning-Based Emergent Communication
by: Boldt, Brendon, et al.
Published: (2024)
by: Boldt, Brendon, et al.
Published: (2024)
Morpheme Induction for Emergent Language
by: Boldt, Brendon, et al.
Published: (2025)
by: Boldt, Brendon, et al.
Published: (2025)
ELCC: the Emergent Language Corpus Collection
by: Boldt, Brendon, et al.
Published: (2024)
by: Boldt, Brendon, et al.
Published: (2024)
Phonotactic Complexity across Dialects
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
by: Yeo, Eunjung, et al.
Published: (2026)
by: Yeo, Eunjung, et al.
Published: (2026)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
by: Bharadwaj, Shikhar, et al.
Published: (2025)
by: Bharadwaj, Shikhar, et al.
Published: (2025)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
by: Bharadwaj, Shikhar, et al.
Published: (2026)
by: Bharadwaj, Shikhar, et al.
Published: (2026)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
by: Yeo, Eunjung
Published: (2024)
by: Yeo, Eunjung
Published: (2024)
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
by: Yeo, Eunjung, et al.
Published: (2025)
by: Yeo, Eunjung, et al.
Published: (2025)
Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
by: Rosero, Karen, et al.
Published: (2025)
by: Rosero, Karen, et al.
Published: (2025)
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
by: Li, Chin-Jou, et al.
Published: (2025)
by: Li, Chin-Jou, et al.
Published: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
by: Sun, Haitong, et al.
Published: (2026)
by: Sun, Haitong, et al.
Published: (2026)
kalvinroberts/RelayModelCode: Relay Model Code
by: Kalvin Roberts
Published: (2026)
by: Kalvin Roberts
Published: (2026)
On-device Streaming Discrete Speech Units
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
Programming by Examples Meets Historical Linguistics: A Large Language Model Based Approach to Sound Law Induction
by: Naik, Atharva, et al.
Published: (2025)
by: Naik, Atharva, et al.
Published: (2025)
Satellite and Mobile Phone Data Reveal How Violence Affects Seasonal Migration in Afghanistan
by: Tai, Xiao Hui, et al.
Published: (2025)
by: Tai, Xiao Hui, et al.
Published: (2025)
MePRiSIA: risk prevention methodology for academic information systems
by: Isabel Cristina Satizábal-Echavarría
Published: (2018)
by: Isabel Cristina Satizábal-Echavarría
Published: (2018)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
by: Hernandez, Abner, et al.
Published: (2026)
by: Hernandez, Abner, et al.
Published: (2026)
OpusLM: A Family of Open Unified Speech Language Models
by: Tian, Jinchuan, et al.
Published: (2025)
by: Tian, Jinchuan, et al.
Published: (2025)
Structure of Quantum Mean Force Gibbs States for Coupled Harmonic Systems
by: Yeo, Joonhyun, et al.
Published: (2024)
by: Yeo, Joonhyun, et al.
Published: (2024)
Role of System-Bath Interaction in Non-Markovian Quantum Brownian Otto Cycles
by: Shim, Haena, et al.
Published: (2026)
by: Shim, Haena, et al.
Published: (2026)
Realizing the phantom-divide crossing with vector and scalar fields
by: Tsujikawa, Shinji
Published: (2026)
by: Tsujikawa, Shinji
Published: (2026)
PRiMeFlow: Capturing Complex Expression Heterogeneity in Perturbation Response Modelling
by: Yan, Zichao, et al.
Published: (2026)
by: Yan, Zichao, et al.
Published: (2026)
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
EmoNews: A Spoken Dialogue System for Expressive News Conversations
by: Matsuura, Ryuki, et al.
Published: (2025)
by: Matsuura, Ryuki, et al.
Published: (2025)
Wav2Gloss: Generating Interlinear Glossed Text from Speech
by: He, Taiqi, et al.
Published: (2024)
by: He, Taiqi, et al.
Published: (2024)
Error Estimation for Adaptive Mesh Refinement in Droplet Simulations
by: Nathawani, Darsh, et al.
Published: (2025)
by: Nathawani, Darsh, et al.
Published: (2025)
A one-dimensional mathematical model for shear-induced droplet formation in co-flowing fluids
by: Nathawani, Darsh, et al.
Published: (2023)
by: Nathawani, Darsh, et al.
Published: (2023)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
by: Yan, Brian, et al.
Published: (2025)
by: Yan, Brian, et al.
Published: (2025)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024)
by: Singh, Harman, et al.
Published: (2024)
Maxiprocessos criminais, corrupção e mídia: uma análise a partir da operação lava jato
by: Raphael Boldt
Published: (2020)
by: Raphael Boldt
Published: (2020)
Similar Items
-
An Empirical Recipe for Universal Phone Recognition
by: Bharadwaj, Shikhar, et al.
Published: (2026) -
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
by: Li, Chin-Jou, et al.
Published: (2025) -
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
by: Choi, Kwanghee, et al.
Published: (2025) -
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
by: Imani, Shima, et al.
Published: (2025) -
XferBench: a Data-Driven Benchmark for Emergent Language
by: Boldt, Brendon, et al.
Published: (2024)