BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
Fuente:
arXiv
Saved in:
| Main Authors: | Reuel, Anka, Hardy, Amelia, Smith, Chandler, Lamparth, Max, Hardy, Malcolm, Kochenderfer, Mykel J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023)
by: Lamparth, Max, et al.
Published: (2023)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024)
by: Hardy, Amelia, et al.
Published: (2024)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
by: Rivera, Juan-Pablo, et al.
Published: (2024)
by: Rivera, Juan-Pablo, et al.
Published: (2024)
Inferring Traffic Models in Terminal Airspace from Flight Tracks and Procedures
by: Jung, Soyeon, et al.
Published: (2023)
by: Jung, Soyeon, et al.
Published: (2023)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
by: Lamparth, Max, et al.
Published: (2026)
by: Lamparth, Max, et al.
Published: (2026)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
by: Haupt, Andreas, et al.
Published: (2026)
by: Haupt, Andreas, et al.
Published: (2026)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
by: Chaubard, Francois, et al.
Published: (2025)
by: Chaubard, Francois, et al.
Published: (2025)
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
by: Fein, Daniel, et al.
Published: (2026)
by: Fein, Daniel, et al.
Published: (2026)
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024)
by: Chaubard, Francois, et al.
Published: (2024)
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
MileBench: Benchmarking MLLMs in Long Context
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
Graph Q-Learning for Combinatorial Optimization
by: Dax, Victoria M., et al.
Published: (2024)
by: Dax, Victoria M., et al.
Published: (2024)
Zono-Conformal Prediction: Zonotope-Based Uncertainty Quantification for Regression and Classification Tasks
by: Lützow, Laura, et al.
Published: (2025)
by: Lützow, Laura, et al.
Published: (2025)
IGL-Bench: Establishing the Comprehensive Benchmark for Imbalanced Graph Learning
by: Qin, Jiawen, et al.
Published: (2024)
by: Qin, Jiawen, et al.
Published: (2024)
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
by: Jafari, Kiana, et al.
Published: (2026)
by: Jafari, Kiana, et al.
Published: (2026)
Imperfect World Models are Exploitable
by: Bhamidipaty, Logan Mondal, et al.
Published: (2026)
by: Bhamidipaty, Logan Mondal, et al.
Published: (2026)
Conditional Deep Generative Models for Belief State Planning
by: Bigeard, Antoine, et al.
Published: (2025)
by: Bigeard, Antoine, et al.
Published: (2025)
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Establishing Best Practices for Building Rigorous Agentic Benchmarks
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Enhanced Importance Sampling through Latent Space Exploration in Normalizing Flows
by: Kruse, Liam A., et al.
Published: (2025)
by: Kruse, Liam A., et al.
Published: (2025)
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
by: Srinivasa, Rakshith S, et al.
Published: (2025)
by: Srinivasa, Rakshith S, et al.
Published: (2025)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024)
by: Hardy, Michael
Published: (2024)
ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors
by: Hardy, Romain, et al.
Published: (2025)
by: Hardy, Romain, et al.
Published: (2025)
Robust Planning for Autonomous Vehicles with Diffusion-Based Failure Samplers
by: Wang, Juanran, et al.
Published: (2025)
by: Wang, Juanran, et al.
Published: (2025)
Measuring Teaching with LLMs
by: Hardy, Michael
Published: (2025)
by: Hardy, Michael
Published: (2025)
A Semi-Decentralized Approach to Multiagent Control
by: Al-Husseini, Mahdi, et al.
Published: (2026)
by: Al-Husseini, Mahdi, et al.
Published: (2026)
Addressing Myopic Constrained POMDP Planning with Recursive Dual Ascent
by: Stocco, Paula, et al.
Published: (2024)
by: Stocco, Paula, et al.
Published: (2024)
Semi-Markovian Planning to Coordinate Aerial and Maritime Medical Evacuation Platforms
by: Al-Husseini, Mahdi, et al.
Published: (2024)
by: Al-Husseini, Mahdi, et al.
Published: (2024)
Optimal Ground Station Selection for Low-Earth Orbiting Satellites
by: Eddy, Duncan, et al.
Published: (2024)
by: Eddy, Duncan, et al.
Published: (2024)
Scene Informer: Anchor-based Occlusion Inference and Trajectory Prediction in Partially Observable Environments
by: Lange, Bernard, et al.
Published: (2023)
by: Lange, Bernard, et al.
Published: (2023)
LeRAAT: LLM-Enabled Real-Time Aviation Advisory Tool
by: Schlichting, Marc R., et al.
Published: (2025)
by: Schlichting, Marc R., et al.
Published: (2025)
SCOUT: A Lightweight Framework for Scenario Coverage Assessment in Autonomous Driving
by: Yildiz, Anil, et al.
Published: (2025)
by: Yildiz, Anil, et al.
Published: (2025)
BetaZero: Belief-State Planning for Long-Horizon POMDPs using Learned Approximations
by: Moss, Robert J., et al.
Published: (2023)
by: Moss, Robert J., et al.
Published: (2023)
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
On Technique Identification and Threat-Actor Attribution using LLMs and Embedding Models
by: Guru, Kyla, et al.
Published: (2025)
by: Guru, Kyla, et al.
Published: (2025)
Similar Items
-
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023) -
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024) -
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
by: Rivera, Juan-Pablo, et al.
Published: (2024) -
Inferring Traffic Models in Terminal Airspace from Flight Tracks and Procedures
by: Jung, Soyeon, et al.
Published: (2023) -
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)