Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phukon, Bornali, Zheng, Xiuwen, Hasegawa-Johnson, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
von: Bae, Jaesung, et al.
Veröffentlicht: (2026)
von: Bae, Jaesung, et al.
Veröffentlicht: (2026)
Aligning Black-box Language Models with Human Judgments
von: Burg, Gerrit J. J. van den, et al.
Veröffentlicht: (2025)
von: Burg, Gerrit J. J. van den, et al.
Veröffentlicht: (2025)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
von: Zheng, Han, et al.
Veröffentlicht: (2026)
von: Zheng, Han, et al.
Veröffentlicht: (2026)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2025)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2025)
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
von: Lee, Wilson Y.
Veröffentlicht: (2026)
von: Lee, Wilson Y.
Veröffentlicht: (2026)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
von: Arakelyan, Erik, et al.
Veröffentlicht: (2024)
von: Arakelyan, Erik, et al.
Veröffentlicht: (2024)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
von: Zhou, Han, et al.
Veröffentlicht: (2024)
von: Zhou, Han, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Representations are More Phonetic than Semantic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
Enhancing Robustness in Biomedical NLI Models: A Probing Approach for Clinical Trials
von: Mustafa, Ata
Veröffentlicht: (2024)
von: Mustafa, Ata
Veröffentlicht: (2024)
VAULT: Vigilant Adversarial Updates via LLM-Driven Retrieval-Augmented Generation for NLI
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
Aligning Large Language Models by On-Policy Self-Judgment
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features
von: Lin, Ye Bhone, et al.
Veröffentlicht: (2025)
von: Lin, Ye Bhone, et al.
Veröffentlicht: (2025)
MICRO: A Lightweight Middleware for Optimizing Cross-store Cross-model Graph-Relation Joins [Technical Report]
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
Inductive Link Prediction in Knowledge Graphs using Path-based Neural Networks
von: Zhang, Canlin, et al.
Veröffentlicht: (2023)
von: Zhang, Canlin, et al.
Veröffentlicht: (2023)
e-Profits: A Business-Aligned Evaluation Metric for Profit-Sensitive Customer Churn Prediction
von: Manzoor, Awais, et al.
Veröffentlicht: (2025)
von: Manzoor, Awais, et al.
Veröffentlicht: (2025)
Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer
von: Qu, Huaizhi, et al.
Veröffentlicht: (2025)
von: Qu, Huaizhi, et al.
Veröffentlicht: (2025)
Mitigating Spurious Correlations in NLI via LLM-Synthesized Counterfactuals and Dynamic Balanced Sampling
von: Jaimes, Christopher Román
Veröffentlicht: (2025)
von: Jaimes, Christopher Román
Veröffentlicht: (2025)
When to Accept Automated Predictions and When to Defer to Human Judgment?
von: Sikar, Daniel, et al.
Veröffentlicht: (2024)
von: Sikar, Daniel, et al.
Veröffentlicht: (2024)
LASER: An LLM-based ASR Scoring and Evaluation Rubric
von: Parulekar, Amruta, et al.
Veröffentlicht: (2025)
von: Parulekar, Amruta, et al.
Veröffentlicht: (2025)
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
von: Zhang, Zheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zheng, et al.
Veröffentlicht: (2026)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
DQE: A Semantic-Aware Evaluation Metric for Time Series Anomaly Detection
von: Li, Yuewei, et al.
Veröffentlicht: (2026)
von: Li, Yuewei, et al.
Veröffentlicht: (2026)
Augmenting Legal Decision Support Systems with LLM-based NLI for Analyzing Social Media Evidence
von: Kadiyala, Ram Mohan Rao, et al.
Veröffentlicht: (2024)
von: Kadiyala, Ram Mohan Rao, et al.
Veröffentlicht: (2024)
Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
Taxon: Hierarchical Tax Code Prediction with Semantically Aligned LLM Expert Guidance
von: Li, Jihang, et al.
Veröffentlicht: (2026)
von: Li, Jihang, et al.
Veröffentlicht: (2026)
Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialogues
von: Hong, Mengze, et al.
Veröffentlicht: (2024)
von: Hong, Mengze, et al.
Veröffentlicht: (2024)
A Theory of Machine Understanding via the Minimum Description Length Principle
von: Zhang, Canlin, et al.
Veröffentlicht: (2025)
von: Zhang, Canlin, et al.
Veröffentlicht: (2025)
Learning Regularities from Data using Spiking Functions: A Theory
von: Zhang, Canlin, et al.
Veröffentlicht: (2024)
von: Zhang, Canlin, et al.
Veröffentlicht: (2024)
Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings
von: Taylor, Russell, et al.
Veröffentlicht: (2025)
von: Taylor, Russell, et al.
Veröffentlicht: (2025)
Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
von: Kumar, Vanya Bannihatti, et al.
Veröffentlicht: (2025)
von: Kumar, Vanya Bannihatti, et al.
Veröffentlicht: (2025)
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
von: Lang, Yicheng, et al.
Veröffentlicht: (2025)
von: Lang, Yicheng, et al.
Veröffentlicht: (2025)
Performance Evaluation of Ising and QUBO Variable Encodings in Boltzmann Machine Learning
von: Hasegawa, Yasushi, et al.
Veröffentlicht: (2025)
von: Hasegawa, Yasushi, et al.
Veröffentlicht: (2025)
Aligning Multiclass Neural Network Classifier Criterion with Task Performance Metrics
von: Li, Deyuan, et al.
Veröffentlicht: (2024)
von: Li, Deyuan, et al.
Veröffentlicht: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
von: Trivedi, Prashant, et al.
Veröffentlicht: (2025)
von: Trivedi, Prashant, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024) -
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026) -
Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
von: Bae, Jaesung, et al.
Veröffentlicht: (2026) -
Aligning Black-box Language Models with Human Judgments
von: Burg, Gerrit J. J. van den, et al.
Veröffentlicht: (2025) -
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)