Evaluating Speech-to-Text Systems with PennSound
Fuente:
arXiv
Saved in:
| Main Authors: | Wright, Jonathan, Liberman, Mark, Ryant, Neville, Fiumara, James |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems
by: Borgholt, Lasse, et al.
Published: (2025)
by: Borgholt, Lasse, et al.
Published: (2025)
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
by: Allbert, Rumi, et al.
Published: (2025)
by: Allbert, Rumi, et al.
Published: (2025)
Conceptors for Semantic Steering
by: Triantafyllopoulos, Ilias, et al.
Published: (2026)
by: Triantafyllopoulos, Ilias, et al.
Published: (2026)
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
by: Gaido, Marco, et al.
Published: (2025)
by: Gaido, Marco, et al.
Published: (2025)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
by: Minixhofer, Christoph, et al.
Published: (2025)
by: Minixhofer, Christoph, et al.
Published: (2025)
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems
by: Iakovenko, Olga, et al.
Published: (2024)
by: Iakovenko, Olga, et al.
Published: (2024)
Evaluating Text Classification Robustness to Part-of-Speech Adversarial Examples
by: Samadi, Anahita, et al.
Published: (2024)
by: Samadi, Anahita, et al.
Published: (2024)
CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
by: Du, Yexing, et al.
Published: (2025)
by: Du, Yexing, et al.
Published: (2025)
Flipping the Dialogue: Training and Evaluating User Language Models
by: Naous, Tarek, et al.
Published: (2025)
by: Naous, Tarek, et al.
Published: (2025)
Comparative Evaluation of Machine Translation Systems on Images with Text
by: Puchol, Blai, et al.
Published: (2026)
by: Puchol, Blai, et al.
Published: (2026)
Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System
by: Du, Yanfan, et al.
Published: (2025)
by: Du, Yanfan, et al.
Published: (2025)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
by: Patel, Fagun, et al.
Published: (2025)
by: Patel, Fagun, et al.
Published: (2025)
MunTTS: A Text-to-Speech System for Mundari
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
Careless Whisper: Speech-to-Text Hallucination Harms
by: Koenecke, Allison, et al.
Published: (2024)
by: Koenecke, Allison, et al.
Published: (2024)
Growing Trees on Sounds: Assessing Strategies for End-to-End Dependency Parsing of Speech
by: Pupier, Adrien, et al.
Published: (2024)
by: Pupier, Adrien, et al.
Published: (2024)
Neural networks for Text-to-Speech evaluation
by: Trofimenko, Ilya, et al.
Published: (2026)
by: Trofimenko, Ilya, et al.
Published: (2026)
The Logovista English-Japanese Machine Translation System
by: Wright, Barton D.
Published: (2026)
by: Wright, Barton D.
Published: (2026)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
by: Sankar, Sanjana, et al.
Published: (2025)
by: Sankar, Sanjana, et al.
Published: (2025)
LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models
by: Khamis, Ahmed Khaled, et al.
Published: (2026)
by: Khamis, Ahmed Khaled, et al.
Published: (2026)
Evaluating the Data Model Robustness of Text-to-SQL Systems Based on Real User Queries
by: Fürst, Jonathan, et al.
Published: (2024)
by: Fürst, Jonathan, et al.
Published: (2024)
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
by: Zhao, Xingjian, et al.
Published: (2025)
by: Zhao, Xingjian, et al.
Published: (2025)
Continuous Speech Tokenizer in Text To Speech
by: Li, Yixing, et al.
Published: (2024)
by: Li, Yixing, et al.
Published: (2024)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
by: Qiang, Chunyu, et al.
Published: (2026)
by: Qiang, Chunyu, et al.
Published: (2026)
A Linguistically Motivated Analysis of Intonational Phrasing in Text-to-Speech Systems: Revealing Gaps in Syntactic Sensitivity
by: Pouw, Charlotte, et al.
Published: (2025)
by: Pouw, Charlotte, et al.
Published: (2025)
Text to Speech System for Meitei Mayek Script
by: Irengbam, Gangular Singh, et al.
Published: (2025)
by: Irengbam, Gangular Singh, et al.
Published: (2025)
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
by: Du, Hongfei, et al.
Published: (2026)
by: Du, Hongfei, et al.
Published: (2026)
Creating an Aligned Corpus of Sound and Text: The Multimodal Corpus of Shakespeare and Milton
by: Agirrezabal, Manex
Published: (2024)
by: Agirrezabal, Manex
Published: (2024)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
by: Alastruey, Belen, et al.
Published: (2023)
by: Alastruey, Belen, et al.
Published: (2023)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
by: Bhogale, Kaushal Santosh, et al.
Published: (2026)
by: Bhogale, Kaushal Santosh, et al.
Published: (2026)
DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
by: Tan, Chao-Hong, et al.
Published: (2025)
by: Tan, Chao-Hong, et al.
Published: (2025)
EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech
by: Hassan, Besher, et al.
Published: (2026)
by: Hassan, Besher, et al.
Published: (2026)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
Whisper-UT: A Unified Translation Framework for Speech and Text
by: Xiao, Cihan, et al.
Published: (2025)
by: Xiao, Cihan, et al.
Published: (2025)
Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text
by: Ruby, Ahmed, et al.
Published: (2026)
by: Ruby, Ahmed, et al.
Published: (2026)
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
by: Du, Yexing, et al.
Published: (2026)
by: Du, Yexing, et al.
Published: (2026)
Cross-lingual Matryoshka Representation Learning across Speech and Text
by: Sy, Yaya, et al.
Published: (2026)
by: Sy, Yaya, et al.
Published: (2026)
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
by: Koshkin, Roman, et al.
Published: (2026)
by: Koshkin, Roman, et al.
Published: (2026)
Similar Items
-
A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems
by: Borgholt, Lasse, et al.
Published: (2025) -
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
by: Allbert, Rumi, et al.
Published: (2025) -
Conceptors for Semantic Steering
by: Triantafyllopoulos, Ilias, et al.
Published: (2026) -
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
by: Gaido, Marco, et al.
Published: (2025) -
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
by: Romero-Díaz, Jacobo, et al.
Published: (2025)