ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sedláček, Šimon, Barahona, Sara, Yusuf, Bolaji, Herrera-Alarcón, Laura, Kesiraju, Santosh, Bolaños, Cecilia, Lozano-Diez, Alicia, Udupa, Sathvik, López, Fernando, Ferner, Allison, Duraiswami, Ramani, Černocký, Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915664700112896
author Sedláček, Šimon
Barahona, Sara
Yusuf, Bolaji
Herrera-Alarcón, Laura
Kesiraju, Santosh
Bolaños, Cecilia
Lozano-Diez, Alicia
Udupa, Sathvik
López, Fernando
Ferner, Allison
Duraiswami, Ramani
Černocký, Jan
author_facet Sedláček, Šimon
Barahona, Sara
Yusuf, Bolaji
Herrera-Alarcón, Laura
Kesiraju, Santosh
Bolaños, Cecilia
Lozano-Diez, Alicia
Udupa, Sathvik
López, Fernando
Ferner, Allison
Duraiswami, Ramani
Černocký, Jan
contents Evaluating open-ended responses from large audio language models (LALMs) is challenging because human annotators often genuinely disagree on answer correctness due to multiple valid interpretations, partial correctness, and subjective judgment. Traditional metrics reporting only mean scores fail to capture this uncertainty. We present ORCA (Open-ended Response Correctness Assessment), a framework that models the variability in human judgments using Beta distributions to predict both expected correctness and uncertainty. Our three-stage annotation framework combines human judgment with structured feedback and iterative refinement to simultaneously curate training data and improve benchmark quality. We collected 11,721 annotations across 3,580 question-answer pairs from 15 LALMs on two audio QA benchmarks, achieving inter-annotator agreement of 0.82 (Krippendorff's alpha). ORCA achieves 0.91 Spearman correlation with mean human judgments, matching or outperforming LLM-judge baselines while providing uncertainty estimates and requiring significantly less compute. We release our models, code, and curated dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09066
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
Sedláček, Šimon
Barahona, Sara
Yusuf, Bolaji
Herrera-Alarcón, Laura
Kesiraju, Santosh
Bolaños, Cecilia
Lozano-Diez, Alicia
Udupa, Sathvik
López, Fernando
Ferner, Allison
Duraiswami, Ramani
Černocký, Jan
Sound
Artificial Intelligence
Computation and Language
Evaluating open-ended responses from large audio language models (LALMs) is challenging because human annotators often genuinely disagree on answer correctness due to multiple valid interpretations, partial correctness, and subjective judgment. Traditional metrics reporting only mean scores fail to capture this uncertainty. We present ORCA (Open-ended Response Correctness Assessment), a framework that models the variability in human judgments using Beta distributions to predict both expected correctness and uncertainty. Our three-stage annotation framework combines human judgment with structured feedback and iterative refinement to simultaneously curate training data and improve benchmark quality. We collected 11,721 annotations across 3,580 question-answer pairs from 15 LALMs on two audio QA benchmarks, achieving inter-annotator agreement of 0.82 (Krippendorff's alpha). ORCA achieves 0.91 Spearman correlation with mean human judgments, matching or outperforming LLM-judge baselines while providing uncertainty estimates and requiring significantly less compute. We release our models, code, and curated dataset.
title ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
topic Sound
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.09066