AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Hoang, Surapaneni, Sidharth, Kalkunte, Akshay, Mehta, Jash, Tiwari, Aman, Bamgbose, Oluwanifemi, Mahajan, Khyati, Shah, Jash, Radhakrishna, Shruthan, Madhusudhan, Sathwik Tejaswi, Yadav, Vikas, Rajeswar, Sai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
by: Hashemi, Masoud, et al.
Published: (2025)
by: Hashemi, Masoud, et al.
Published: (2025)
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
by: Maheshwary, Rishabh, et al.
Published: (2024)
by: Maheshwary, Rishabh, et al.
Published: (2024)
Apriel-1.5-15b-Thinker
by: Radhakrishna, Shruthan, et al.
Published: (2025)
by: Radhakrishna, Shruthan, et al.
Published: (2025)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025)
by: Maheshwary, Rishabh, et al.
Published: (2025)
Apriel-Nemotron-15B-Thinker
by: Radhakrishna, Shruthan, et al.
Published: (2025)
by: Radhakrishna, Shruthan, et al.
Published: (2025)
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
by: Madhusudhan, Nishanth, et al.
Published: (2024)
by: Madhusudhan, Nishanth, et al.
Published: (2024)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
by: Tiwari, Aman, et al.
Published: (2024)
by: Tiwari, Aman, et al.
Published: (2024)
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
by: Rajeev, Meghana, et al.
Published: (2025)
by: Rajeev, Meghana, et al.
Published: (2025)
DeepSRGM -- Sequence Classification and Ranking in Indian Classical Music with Deep Learning
by: Madhusudhan, Sathwik Tejaswi, et al.
Published: (2024)
by: Madhusudhan, Sathwik Tejaswi, et al.
Published: (2024)
Grammar Search for Multi-Agent Systems
by: Singh, Mayank, et al.
Published: (2025)
by: Singh, Mayank, et al.
Published: (2025)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
by: Pattnaik, Pulkit, et al.
Published: (2024)
by: Pattnaik, Pulkit, et al.
Published: (2024)
Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
by: Etzine, Bryan, et al.
Published: (2025)
by: Etzine, Bryan, et al.
Published: (2025)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Evidence for Complement Activation in Preeclampsia Placenta and Its Presence in Circulation
by: Shibin Cheng, et al.
Published: (2025)
by: Shibin Cheng, et al.
Published: (2025)
Spatial Competence Benchmark
by: Vira, Jash, et al.
Published: (2026)
by: Vira, Jash, et al.
Published: (2026)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
Apriel-H1: Towards Efficient Enterprise Reasoning Models
by: Ostapenko, Oleksiy, et al.
Published: (2025)
by: Ostapenko, Oleksiy, et al.
Published: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
Multi-Class Boundary Extraction from Implicit Representations
by: Vira, Jash, et al.
Published: (2026)
by: Vira, Jash, et al.
Published: (2026)
Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience
by: Sayana, Krishna, et al.
Published: (2026)
by: Sayana, Krishna, et al.
Published: (2026)
Practical Guide for Causal Pathways and Sub-group Disparity Analysis
by: Kohankhaki, Farnaz, et al.
Published: (2024)
by: Kohankhaki, Farnaz, et al.
Published: (2024)
Structure-Augmented Reasoning Generation
by: Parekh, Jash Rajesh, et al.
Published: (2025)
by: Parekh, Jash Rajesh, et al.
Published: (2025)
Suppression of RBC alloimmunization and regulation of CD4 + T cell dependence by C3 is not due to genetic confounders in mice
by: Arijita Jash, et al.
Published: (2025)
by: Arijita Jash, et al.
Published: (2025)
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
by: Bogavelli, Tara, et al.
Published: (2026)
by: Bogavelli, Tara, et al.
Published: (2026)
mmWave Radar Aware Dual-Conditioned GAN for Speech Reconstruction of Signals With Low SNR
by: Karani, Jash, et al.
Published: (2026)
by: Karani, Jash, et al.
Published: (2026)
Homogenization Effects of Large Language Models on Human Creative Ideation
by: Anderson, Barrett R., et al.
Published: (2024)
by: Anderson, Barrett R., et al.
Published: (2024)
Distilling Aggregated Knowledge for Weakly-Supervised Video Anomaly Detection
by: Dalvi, Jash, et al.
Published: (2024)
by: Dalvi, Jash, et al.
Published: (2024)
Vision Transformers for Cosmological Fields: Application to Weak Lensing Mass Maps
by: Kakadia, Jash, et al.
Published: (2025)
by: Kakadia, Jash, et al.
Published: (2025)
Endometrial Cancer Diagnosis During IVF Treatment With Subsequent Live Birth Following Fertility‐Sparing Approach: A Case Report
by: Jash Bhankharia, et al.
Published: (2026)
by: Jash Bhankharia, et al.
Published: (2026)
Discharge quenching mechanism and performance of RPWELL with tunable 3D printed resistive plates, charge evacuation in semiconductive glass RPWELL and discharge quenching for Cryogenic-RWELL over a wide range of resistivity
by: Jash, Abhik, et al.
Published: (2024)
by: Jash, Abhik, et al.
Published: (2024)
Prompting with Phonemes: Enhancing LLMs' Multilinguality for Non-Latin Script Languages
by: Nguyen, Hoang H, et al.
Published: (2024)
by: Nguyen, Hoang H, et al.
Published: (2024)
Trashbusters: Deep Learning Approach for Litter Detection and Tracking
by: Jain, Kashish, et al.
Published: (2024)
by: Jain, Kashish, et al.
Published: (2024)
Taub-NUT Instanton as the Self-dual Analog of Kerr
by: Desai, Jash, et al.
Published: (2024)
by: Desai, Jash, et al.
Published: (2024)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
FakeWatch: A Framework for Detecting Fake News to Ensure Credible Elections
by: Raza, Shaina, et al.
Published: (2024)
by: Raza, Shaina, et al.
Published: (2024)
Unlocking Bias Detection: Leveraging Transformer-Based Models for Content Analysis
by: Raza, Shaina, et al.
Published: (2023)
by: Raza, Shaina, et al.
Published: (2023)
User Embedding Model for Personalized Language Prompting
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Metal-organic Frameworks in Semiconductor Devices: A Revised Version
by: Parashar, Ranjeev Kumar, et al.
Published: (2024)
by: Parashar, Ranjeev Kumar, et al.
Published: (2024)
Similar Items
-
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
by: Hashemi, Masoud, et al.
Published: (2025) -
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
by: Maheshwary, Rishabh, et al.
Published: (2024) -
Apriel-1.5-15b-Thinker
by: Radhakrishna, Shruthan, et al.
Published: (2025) -
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025) -
Apriel-Nemotron-15B-Thinker
by: Radhakrishna, Shruthan, et al.
Published: (2025)