AHELM: A Holistic Evaluation of Audio-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Tony, Tu, Haoqin, Wong, Chi Heem, Wang, Zijun, Yang, Siwei, Mai, Yifan, Zhou, Yuyin, Xie, Cihang, Liang, Percy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VHELM: A Holistic Evaluation of Vision Language Models
von: Lee, Tony, et al.
Veröffentlicht: (2024)
von: Lee, Tony, et al.
Veröffentlicht: (2024)
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
von: Roberts, Josselin Somerville, et al.
Veröffentlicht: (2024)
von: Roberts, Josselin Somerville, et al.
Veröffentlicht: (2024)
A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
von: Chen, Hardy, et al.
Veröffentlicht: (2025)
von: Chen, Hardy, et al.
Veröffentlicht: (2025)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
von: Wang, Zijun, et al.
Veröffentlicht: (2024)
von: Wang, Zijun, et al.
Veröffentlicht: (2024)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
von: Chen, Hardy, et al.
Veröffentlicht: (2026)
von: Chen, Hardy, et al.
Veröffentlicht: (2026)
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
von: Wu, Juncheng, et al.
Veröffentlicht: (2025)
von: Wu, Juncheng, et al.
Veröffentlicht: (2025)
ViLBench: A Suite for Vision-Language Process Reward Modeling
von: Tu, Haoqin, et al.
Veröffentlicht: (2025)
von: Tu, Haoqin, et al.
Veröffentlicht: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline
von: Lee, Tony, et al.
Veröffentlicht: (2026)
von: Lee, Tony, et al.
Veröffentlicht: (2026)
Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
von: Han, Qijun, et al.
Veröffentlicht: (2026)
von: Han, Qijun, et al.
Veröffentlicht: (2026)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
von: Lu, Ke-Han, et al.
Veröffentlicht: (2026)
von: Lu, Ke-Han, et al.
Veröffentlicht: (2026)
SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
von: Susanto, Yosephine, et al.
Veröffentlicht: (2025)
von: Susanto, Yosephine, et al.
Veröffentlicht: (2025)
Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
von: Bedi, Suhana, et al.
Veröffentlicht: (2025)
von: Bedi, Suhana, et al.
Veröffentlicht: (2025)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
Reliable and Efficient Amortized Model-based Evaluation
von: Truong, Sang, et al.
Veröffentlicht: (2025)
von: Truong, Sang, et al.
Veröffentlicht: (2025)
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
von: Qiu, Shi, et al.
Veröffentlicht: (2025)
von: Qiu, Shi, et al.
Veröffentlicht: (2025)
Evaluating Human-Language Model Interaction
von: Lee, Mina, et al.
Veröffentlicht: (2022)
von: Lee, Mina, et al.
Veröffentlicht: (2022)
Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Study
von: Hou, Guanyu, et al.
Veröffentlicht: (2025)
von: Hou, Guanyu, et al.
Veröffentlicht: (2025)
Pardon? Evaluating Conversational Repair in Large Audio-Language Models
von: Huang, Shuanghong, et al.
Veröffentlicht: (2026)
von: Huang, Shuanghong, et al.
Veröffentlicht: (2026)
Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
On the Entropy Calibration of Language Models
von: Cao, Steven, et al.
Veröffentlicht: (2025)
von: Cao, Steven, et al.
Veröffentlicht: (2025)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
Holistic Evaluations of Topic Models
von: Compton, Thomas
Veröffentlicht: (2025)
von: Compton, Thomas
Veröffentlicht: (2025)
CHAI: Command Hijacking against embodied AI
von: Burbano, Luis, et al.
Veröffentlicht: (2025)
von: Burbano, Luis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VHELM: A Holistic Evaluation of Vision Language Models
von: Lee, Tony, et al.
Veröffentlicht: (2024) -
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
von: Wu, Juncheng, et al.
Veröffentlicht: (2026) -
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
von: Roberts, Josselin Somerville, et al.
Veröffentlicht: (2024) -
A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?
von: Xie, Yunfei, et al.
Veröffentlicht: (2024) -
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
von: Chen, Hardy, et al.
Veröffentlicht: (2025)