A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Borgholt, Lasse, Havtorn, Jakob, Igel, Christian, Maaløe, Lars, Tan, Zheng-Hua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces
by: Nix, Arne, et al.
Published: (2026)
by: Nix, Arne, et al.
Published: (2026)
An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare Records
by: Edin, Joakim, et al.
Published: (2024)
by: Edin, Joakim, et al.
Published: (2024)
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems
by: Iakovenko, Olga, et al.
Published: (2024)
by: Iakovenko, Olga, et al.
Published: (2024)
Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text
by: Ruby, Ahmed, et al.
Published: (2026)
by: Ruby, Ahmed, et al.
Published: (2026)
Evaluating Speech-to-Text Systems with PennSound
by: Wright, Jonathan, et al.
Published: (2025)
by: Wright, Jonathan, et al.
Published: (2025)
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
by: Allbert, Rumi, et al.
Published: (2025)
by: Allbert, Rumi, et al.
Published: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
by: Koshkin, Roman, et al.
Published: (2026)
by: Koshkin, Roman, et al.
Published: (2026)
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
by: Du, Yexing, et al.
Published: (2024)
by: Du, Yexing, et al.
Published: (2024)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
by: Pareras, Oriol, et al.
Published: (2025)
by: Pareras, Oriol, et al.
Published: (2025)
Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2024)
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2024)
Text-Utilization for Encoder-dominated Speech Recognition Models
by: Zeyer, Albert, et al.
Published: (2026)
by: Zeyer, Albert, et al.
Published: (2026)
Towards Better Text-to-Image Generation Alignment via Attention Modulation
by: Wu, Yihang, et al.
Published: (2024)
by: Wu, Yihang, et al.
Published: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
by: Liu, Henglyu, et al.
Published: (2025)
by: Liu, Henglyu, et al.
Published: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
by: Ma, Ziyang, et al.
Published: (2023)
by: Ma, Ziyang, et al.
Published: (2023)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
by: Zhang, Pei, et al.
Published: (2025)
by: Zhang, Pei, et al.
Published: (2025)
CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
by: Du, Yexing, et al.
Published: (2025)
by: Du, Yexing, et al.
Published: (2025)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
by: Ye, Zongli, et al.
Published: (2025)
by: Ye, Zongli, et al.
Published: (2025)
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
by: Gaido, Marco, et al.
Published: (2025)
by: Gaido, Marco, et al.
Published: (2025)
Improvement in Sign Language Translation Using Text CTC Alignment
by: Tan, Sihan, et al.
Published: (2024)
by: Tan, Sihan, et al.
Published: (2024)
Modeling Text-Label Alignment for Hierarchical Text Classification
by: Kumar, Ashish, et al.
Published: (2024)
by: Kumar, Ashish, et al.
Published: (2024)
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models
by: Xu, Wenze, et al.
Published: (2025)
by: Xu, Wenze, et al.
Published: (2025)
Learning Personalized Alignment for Evaluating Open-ended Text Generation
by: Wang, Danqing, et al.
Published: (2023)
by: Wang, Danqing, et al.
Published: (2023)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
by: Lu, Ke-Han, et al.
Published: (2024)
by: Lu, Ke-Han, et al.
Published: (2024)
Are Clinical T5 Models Better for Clinical Text?
by: Li, Yahan, et al.
Published: (2024)
by: Li, Yahan, et al.
Published: (2024)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
by: Cornell, Samuele, et al.
Published: (2024)
by: Cornell, Samuele, et al.
Published: (2024)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
by: Wang, Guansu, et al.
Published: (2025)
by: Wang, Guansu, et al.
Published: (2025)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems
by: Meng, Qingliang, et al.
Published: (2025)
by: Meng, Qingliang, et al.
Published: (2025)
Revisiting N-Gram Models: Their Impact in Modern Neural Networks for Handwritten Text Recognition
by: Tarride, Solène, et al.
Published: (2024)
by: Tarride, Solène, et al.
Published: (2024)
Modern Models, Medieval Texts: A POS Tagging Study of Old Occitan
by: Schöffel, Matthias, et al.
Published: (2025)
by: Schöffel, Matthias, et al.
Published: (2025)
Computational Measurement of Political Positions: A Review of Text-Based Ideal Point Estimation Algorithms
by: Parschan, Patrick, et al.
Published: (2025)
by: Parschan, Patrick, et al.
Published: (2025)
Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models
by: Xiang, Bajian, et al.
Published: (2025)
by: Xiang, Bajian, et al.
Published: (2025)
Enhancing Multilingual Voice Toxicity Detection with Speech-Text Alignment
by: Liu, Joseph, et al.
Published: (2024)
by: Liu, Joseph, et al.
Published: (2024)
Similar Items
-
Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces
by: Nix, Arne, et al.
Published: (2026) -
An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare Records
by: Edin, Joakim, et al.
Published: (2024) -
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems
by: Iakovenko, Olga, et al.
Published: (2024) -
Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text
by: Ruby, Ahmed, et al.
Published: (2026) -
Evaluating Speech-to-Text Systems with PennSound
by: Wright, Jonathan, et al.
Published: (2025)