MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Guoli, Bai, Haoping, Ma, Shuang, Nan, Feng, Sun, Yanchao, Xu, Zhaoyang, Ma, Shen, Lu, Jiarui, Kong, Xiang, Zhang, Aonan, Yap, Dian Ang, zhang, Yizhe, Ahnert, Karsten, Kamath, Vik, Berglund, Mathias, Walsh, Dominic, Gindele, Tobias, Wiest, Juergen, Lai, Zhengfeng, Wang, Xiaoming, Shan, Jiulong, Cao, Meng, Pang, Ruoming, Wang, Zirui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
von: Lu, Jiarui, et al.
Veröffentlicht: (2024)
von: Lu, Jiarui, et al.
Veröffentlicht: (2024)
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
BERT Learns (and Teaches) Chemistry
von: Payne, Josh, et al.
Veröffentlicht: (2020)
von: Payne, Josh, et al.
Veröffentlicht: (2020)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024)
von: Sakshi, S, et al.
Veröffentlicht: (2024)
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo
von: Feng, Shengyu, et al.
Veröffentlicht: (2024)
von: Feng, Shengyu, et al.
Veröffentlicht: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
Legendre's Conjecture: A Proof via the Tower Sieve with Multi-Sieve Compensation Method
von: zhang, xiang, et al.
Veröffentlicht: (2026)
von: zhang, xiang, et al.
Veröffentlicht: (2026)
Bogomol'nyi Equations in Mixed Product Chern-Simons Theories Governing Charged Vortices and Antivortices
von: Xu, Aonan
Veröffentlicht: (2026)
von: Xu, Aonan
Veröffentlicht: (2026)
Synthetic bootstrapped pretraining
von: Yang, Zitong, et al.
Veröffentlicht: (2025)
von: Yang, Zitong, et al.
Veröffentlicht: (2025)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
Adversarial Feature Disentanglement for Bias-Invariant Prediction of Zigong Lantern User Preferences
von: haoyang, zhang, et al.
Veröffentlicht: (2025)
von: haoyang, zhang, et al.
Veröffentlicht: (2025)
Python code for modeling suitable habitats of tea (Camellia sinensis) under future climate scenarios in China
von: zhang
Veröffentlicht: (2025)
von: zhang
Veröffentlicht: (2025)
Detection of protein symmetry and structural rearrangements using secondary structure elements
von: Runfeng Lin, et al.
Veröffentlicht: (2026)
von: Runfeng Lin, et al.
Veröffentlicht: (2026)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
Multilevel Graph Reinforcement Learning for Consistent Cognitive Decision-making in Heterogeneous Mixed Autonomy
von: Gao, Xin, et al.
Veröffentlicht: (2024)
von: Gao, Xin, et al.
Veröffentlicht: (2024)
On irreducible germs of generic morphisms
von: Kulikov, Vik. S.
Veröffentlicht: (2025)
von: Kulikov, Vik. S.
Veröffentlicht: (2025)
The Dilemma of Carbon‐Conscious Consumers: A Multi‐Study Investigation of Carbon Transparency in AI Use
von: Vik Naidoo
Veröffentlicht: (2026)
von: Vik Naidoo
Veröffentlicht: (2026)
AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image Editing
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2023)
Resources consulted and useful for Information Literacy
von: Wiest, Natalie H.
Veröffentlicht: ()
von: Wiest, Natalie H.
Veröffentlicht: ()
Information Literacy: Experiences at Texas A&M University at Galveston
von: Wiest, Natalie H.
Veröffentlicht: ()
von: Wiest, Natalie H.
Veröffentlicht: ()
Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses
von: Rupprecht, Jens, et al.
Veröffentlicht: (2025)
von: Rupprecht, Jens, et al.
Veröffentlicht: (2025)
Homologous nodes in annotated complex networks
von: Moon, Sung Soo, et al.
Veröffentlicht: (2025)
von: Moon, Sung Soo, et al.
Veröffentlicht: (2025)
Estranged Predictions: Measuring Semantic Category Disruption with Masked Language Modelling
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
Fair Pairs: Fairness-Aware Ranking Recovery from Pairwise Comparisons
von: Ahnert, Georg, et al.
Veröffentlicht: (2024)
von: Ahnert, Georg, et al.
Veröffentlicht: (2024)
AlloDesigner
von: zhang, yangyang
Veröffentlicht: (2026)
von: zhang, yangyang
Veröffentlicht: (2026)
Rigid Constraint Self-Pressurized Spherical Fusion Device Research
von: zhang, qin
Veröffentlicht: (2026)
von: zhang, qin
Veröffentlicht: (2026)
Supplementary data
von: hanghang, zhang
Veröffentlicht: (2025)
von: hanghang, zhang
Veröffentlicht: (2025)
超流以太统一理论框架——从经典连续介质力学到量子现象的统一
von: zhang, yongxian
Veröffentlicht: (2026)
von: zhang, yongxian
Veröffentlicht: (2026)
IMS-Data
von: zhang, shuo
Veröffentlicht: (2025)
von: zhang, shuo
Veröffentlicht: (2025)
Pre-decoherence Matrix Product State Simulation for Quantum Systems: Entanglement Control and Compression Efficiency
von: zhang, xiaokun
Veröffentlicht: (2025)
von: zhang, xiaokun
Veröffentlicht: (2025)
test data and result files of PSOSP (Prophage SOS-dependency Predictor)
von: zhang, mujie
Veröffentlicht: (2025)
von: zhang, mujie
Veröffentlicht: (2025)
Timeless Mathematics
von: zhang, dong
Veröffentlicht: (2026)
von: zhang, dong
Veröffentlicht: (2026)
SmartCIF processing log
von: zhang, qixiang
Veröffentlicht: (2026)
von: zhang, qixiang
Veröffentlicht: (2026)
Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
Big AI Models for 6G Wireless Networks: Opportunities, Challenges, and Research Directions
von: Chen, Zirui, et al.
Veröffentlicht: (2023)
von: Chen, Zirui, et al.
Veröffentlicht: (2023)
A new species of Pauroaspis Tang (Hemiptera: Coccomorpha: Asterolecaniidae) from southwestern China, with a key to species worldwide
von: Li, Yu-Ang, et al.
Veröffentlicht: (2025)
von: Li, Yu-Ang, et al.
Veröffentlicht: (2025)
Processus aléatoires et applications -- Algorithmes MCMC et vitesse de convergence
von: Berglund, Nils
Veröffentlicht: (2024)
von: Berglund, Nils
Veröffentlicht: (2024)
Ähnliche Einträge
-
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
von: Lu, Jiarui, et al.
Veröffentlicht: (2024) -
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
von: Kumar, Sonal, et al.
Veröffentlicht: (2025) -
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
von: Du, Xianzhi, et al.
Veröffentlicht: (2024) -
BERT Learns (and Teaches) Chemistry
von: Payne, Josh, et al.
Veröffentlicht: (2020) -
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024)