ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Sedláček, Šimon, Barahona, Sara, Yusuf, Bolaji, Herrera-Alarcón, Laura, Kesiraju, Santosh, Bolaños, Cecilia, Lozano-Diez, Alicia, Udupa, Sathvik, López, Fernando, Ferner, Allison, Duraiswami, Ramani, Černocký, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
by: Vendrame, Katia, et al.
Published: (2025)
by: Vendrame, Katia, et al.
Published: (2025)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
by: Sedláček, Šimon, et al.
Published: (2025)
by: Sedláček, Šimon, et al.
Published: (2025)
FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
by: Kesiraju, Santosh, et al.
Published: (2026)
by: Kesiraju, Santosh, et al.
Published: (2026)
Aligning Pre-trained Models for Spoken Language Translation
by: Sedláček, Šimon, et al.
Published: (2024)
by: Sedláček, Šimon, et al.
Published: (2024)
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
by: Hegde, Pradyoth, et al.
Published: (2025)
by: Hegde, Pradyoth, et al.
Published: (2025)
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
by: Polok, Alexander, et al.
Published: (2025)
by: Polok, Alexander, et al.
Published: (2025)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
by: Udupa, Sathvik, et al.
Published: (2025)
by: Udupa, Sathvik, et al.
Published: (2025)
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
by: Kumar, Sonal, et al.
Published: (2025)
by: Kumar, Sonal, et al.
Published: (2025)
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
by: Yusuf, Bolaji, et al.
Published: (2024)
by: Yusuf, Bolaji, et al.
Published: (2024)
Room Impulse Response Synthesis via Differentiable Feedback Delay Networks for Efficient Spatial Audio Rendering
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
by: Pešán, Jan, et al.
Published: (2024)
by: Pešán, Jan, et al.
Published: (2024)
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
by: Polok, Alexander, et al.
Published: (2024)
by: Polok, Alexander, et al.
Published: (2024)
Analytical Exploration of Spatial Audio Cues: A Differentiable Multi-Sphere Scattering Model
by: Galougah, Siminfar Samakoush, et al.
Published: (2026)
by: Galougah, Siminfar Samakoush, et al.
Published: (2026)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
Auditing Algorithmic Bias in Transformer-Based Trading
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
ViscoReg: Neural Signed Distance Functions via Viscosity Solutions
by: Krishnan, Meenakshi, et al.
Published: (2025)
by: Krishnan, Meenakshi, et al.
Published: (2025)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
by: Anand, Nishit, et al.
Published: (2024)
by: Anand, Nishit, et al.
Published: (2024)
A Bayesian Multilingual Document Model for Zero-shot Topic Identification and Discovery
by: Kesiraju, Santosh, et al.
Published: (2020)
by: Kesiraju, Santosh, et al.
Published: (2020)
Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
by: Galougah, Siminfar Samakoush, et al.
Published: (2024)
by: Galougah, Siminfar Samakoush, et al.
Published: (2024)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
by: Yang, Chao-Han Huck, et al.
Published: (2025)
by: Yang, Chao-Han Huck, et al.
Published: (2025)
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
by: Lassoued, Aymen, et al.
Published: (2026)
by: Lassoued, Aymen, et al.
Published: (2026)
Biomimetic Frontend for Differentiable Audio Processing
by: Famularo, Ruolan Leslie, et al.
Published: (2024)
by: Famularo, Ruolan Leslie, et al.
Published: (2024)
Discovering phoneme-specific critical articulators through a data-driven approach
by: Bandekar, Jesuraj, et al.
Published: (2025)
by: Bandekar, Jesuraj, et al.
Published: (2025)
RECAP: Retrieval-Augmented Audio Captioning
by: Ghosh, Sreyan, et al.
Published: (2023)
by: Ghosh, Sreyan, et al.
Published: (2023)
Quantifying Document Impact in RAG-LLMs
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
On The Application of Linear Attention in Multimodal Transformers
by: Gerami, Armin, et al.
Published: (2026)
by: Gerami, Armin, et al.
Published: (2026)
3D Gaussian Splatting with Normal Information for Mesh Extraction and Improved Rendering
by: Krishnan, Meenakshi, et al.
Published: (2025)
by: Krishnan, Meenakshi, et al.
Published: (2025)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
Written Term Detection Improves Spoken Term Detection
by: Yusuf, Bolaji, et al.
Published: (2024)
by: Yusuf, Bolaji, et al.
Published: (2024)
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
by: Sakshi, S, et al.
Published: (2025)
by: Sakshi, S, et al.
Published: (2025)
Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models
by: Yusuf, Bolaji, et al.
Published: (2024)
by: Yusuf, Bolaji, et al.
Published: (2024)
Robustness assessment of large audio language models in multiple-choice evaluation
by: López, Fernando, et al.
Published: (2025)
by: López, Fernando, et al.
Published: (2025)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
FAST: Factorizable Attention for Speeding up Transformers
by: Gerami, Armin, et al.
Published: (2024)
by: Gerami, Armin, et al.
Published: (2024)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
Morphometric analysis of postnatal lung development in the gray short‐tailed opossum ( Monodelphis domestica ): An ultrastructural study
by: Kirsten Ferner
Published: (2025)
by: Kirsten Ferner
Published: (2025)
Development of the pulmonary vasculature in the gray short‐tailed opossum ( Monodelphis domestica )— 3D reconstruction by microcomputed tomography
by: Kirsten Ferner
Published: (2024)
by: Kirsten Ferner
Published: (2024)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
by: Sakshi, S, et al.
Published: (2024)
by: Sakshi, S, et al.
Published: (2024)
Similar Items
-
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
by: Vendrame, Katia, et al.
Published: (2025) -
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
by: Sedláček, Šimon, et al.
Published: (2025) -
FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
by: Kesiraju, Santosh, et al.
Published: (2026) -
Aligning Pre-trained Models for Spoken Language Translation
by: Sedláček, Šimon, et al.
Published: (2024) -
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
by: Hegde, Pradyoth, et al.
Published: (2025)