Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hofmann, Aris, Vejsbjerg, Inge, Salwala, Dhaval, Daly, Elizabeth M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Humble AI in the real-world: the case of algorithmic hiring
von: Nair, Rahul, et al.
Veröffentlicht: (2025)
von: Nair, Rahul, et al.
Veröffentlicht: (2025)
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models
von: Tirupathi, Seshu, et al.
Veröffentlicht: (2025)
von: Tirupathi, Seshu, et al.
Veröffentlicht: (2025)
A Multi-Layered Research Framework for Human-Centered AI: Defining the Path to Explainability and Trust
von: De Silva, Chameera, et al.
Veröffentlicht: (2025)
von: De Silva, Chameera, et al.
Veröffentlicht: (2025)
SPHERE: An Evaluation Card for Human-AI Systems
von: Ma, Qianou, et al.
Veröffentlicht: (2025)
von: Ma, Qianou, et al.
Veröffentlicht: (2025)
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
von: Sokol, Anna, et al.
Veröffentlicht: (2024)
von: Sokol, Anna, et al.
Veröffentlicht: (2024)
The AI-DEC: A Card-based Design Method for User-centered AI Explanations
von: Lee, Christine P, et al.
Veröffentlicht: (2024)
von: Lee, Christine P, et al.
Veröffentlicht: (2024)
MHDash: An Online Platform for Benchmarking Mental Health-Aware AI Assistants
von: Zhang, Yihe, et al.
Veröffentlicht: (2026)
von: Zhang, Yihe, et al.
Veröffentlicht: (2026)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
von: Henry, Felix, et al.
Veröffentlicht: (2026)
von: Henry, Felix, et al.
Veröffentlicht: (2026)
GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety
von: Fan, Changxuan, et al.
Veröffentlicht: (2026)
von: Fan, Changxuan, et al.
Veröffentlicht: (2026)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
From LLM-Driven Trading Card Generation to Procedural Relatedness: A Pokémon Case Study
von: Pfau, Johannes, et al.
Veröffentlicht: (2026)
von: Pfau, Johannes, et al.
Veröffentlicht: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
von: Wang, Qiaosi, et al.
Veröffentlicht: (2025)
von: Wang, Qiaosi, et al.
Veröffentlicht: (2025)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
von: Yang, Qinglong, et al.
Veröffentlicht: (2025)
von: Yang, Qinglong, et al.
Veröffentlicht: (2025)
LibEER: A Comprehensive Benchmark and Algorithm Library for EEG-based Emotion Recognition
von: Liu, Huan, et al.
Veröffentlicht: (2024)
von: Liu, Huan, et al.
Veröffentlicht: (2024)
A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI
von: Bhatt, Deep, et al.
Veröffentlicht: (2024)
von: Bhatt, Deep, et al.
Veröffentlicht: (2024)
IDRBench: Interactive Deep Research Benchmark
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2026)
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2026)
Benchmarking LLM Tool-Use in the Wild
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
AutoML in The Wild: Obstacles, Workarounds, and Expectations
von: Sun, Yuan, et al.
Veröffentlicht: (2023)
von: Sun, Yuan, et al.
Veröffentlicht: (2023)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
von: Luera, Reuben A., et al.
Veröffentlicht: (2025)
von: Luera, Reuben A., et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
von: Kim, Yoonsu, et al.
Veröffentlicht: (2025)
von: Kim, Yoonsu, et al.
Veröffentlicht: (2025)
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning
von: Lee, Dong Won, et al.
Veröffentlicht: (2025)
von: Lee, Dong Won, et al.
Veröffentlicht: (2025)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
von: Zhu, Ming, et al.
Veröffentlicht: (2026)
von: Zhu, Ming, et al.
Veröffentlicht: (2026)
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
von: Bagehorn, Frank, et al.
Veröffentlicht: (2025)
von: Bagehorn, Frank, et al.
Veröffentlicht: (2025)
Evaluating AI Alignment in LLMs: Output Analysis of Value Priorities Across 75 Models with Human Benchmarking
von: Lau, Gabriel Rongyang, et al.
Veröffentlicht: (2025)
von: Lau, Gabriel Rongyang, et al.
Veröffentlicht: (2025)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
von: Gong, Yichen, et al.
Veröffentlicht: (2026)
von: Gong, Yichen, et al.
Veröffentlicht: (2026)
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
von: Li, Charlotte, et al.
Veröffentlicht: (2025)
von: Li, Charlotte, et al.
Veröffentlicht: (2025)
The Role of AI in Peer Support for Young People: A Study of Preferences for Human- and AI-Generated Responses
von: Young, Jordyn, et al.
Veröffentlicht: (2024)
von: Young, Jordyn, et al.
Veröffentlicht: (2024)
Benchmark It Yourself (BIY): Preparing a Dataset and Benchmarking AI Models for Scatterplot-Related Tasks
von: Palmeiro, João, et al.
Veröffentlicht: (2025)
von: Palmeiro, João, et al.
Veröffentlicht: (2025)
Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots
von: Zhang, Shiquan, et al.
Veröffentlicht: (2026)
von: Zhang, Shiquan, et al.
Veröffentlicht: (2026)
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Creativity Benchmark: A benchmark for marketing creativity for large language models
von: Bhat, Ninad, et al.
Veröffentlicht: (2025)
von: Bhat, Ninad, et al.
Veröffentlicht: (2025)
Human-Centered Automation
von: Toxtli, Carlos
Veröffentlicht: (2024)
von: Toxtli, Carlos
Veröffentlicht: (2024)
SECURE: Benchmarking Large Language Models for Cybersecurity
von: Bhusal, Dipkamal, et al.
Veröffentlicht: (2024)
von: Bhusal, Dipkamal, et al.
Veröffentlicht: (2024)
AutoS$^2$earch: Unlocking the Reasoning Potential of Large Models for Web-based Source Search
von: Zhu, Zhengqiu, et al.
Veröffentlicht: (2025)
von: Zhu, Zhengqiu, et al.
Veröffentlicht: (2025)
AutoGameUI: Constructing High-Fidelity GameUI via Multimodal Correspondence Matching
von: Tang, Zhongliang, et al.
Veröffentlicht: (2024)
von: Tang, Zhongliang, et al.
Veröffentlicht: (2024)
Automated Visualization Makeovers with LLMs
von: Gangwar, Siddharth, et al.
Veröffentlicht: (2025)
von: Gangwar, Siddharth, et al.
Veröffentlicht: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Humble AI in the real-world: the case of algorithmic hiring
von: Nair, Rahul, et al.
Veröffentlicht: (2025) -
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models
von: Tirupathi, Seshu, et al.
Veröffentlicht: (2025) -
A Multi-Layered Research Framework for Human-Centered AI: Defining the Path to Explainability and Trust
von: De Silva, Chameera, et al.
Veröffentlicht: (2025) -
SPHERE: An Evaluation Card for Human-AI Systems
von: Ma, Qianou, et al.
Veröffentlicht: (2025) -
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
von: Sokol, Anna, et al.
Veröffentlicht: (2024)