Fantastic Bugs and Where to Find Them in AI Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Truong, Sang, Tu, Yuheng, Hardy, Michael, Reuel, Anka, Tang, Zeyu, Burapacheep, Jirayu, Perera, Jonathan, Uwakwe, Chibuike, Domingue, Ben, Haber, Nick, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
by: Haupt, Andreas, et al.
Published: (2026)
by: Haupt, Andreas, et al.
Published: (2026)
Your Classifier Can Be Secretly a Likelihood-Based OOD Detector
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
Fantastic Pretraining Optimizers and Where to Find Them
by: Wen, Kaiyue, et al.
Published: (2025)
by: Wen, Kaiyue, et al.
Published: (2025)
Fantastic Biases (What are They) and Where to Find Them
by: Barriere, Valentin
Published: (2024)
by: Barriere, Valentin
Published: (2024)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
TurboVSR: Fantastic Video Upscalers and Where to Find Them
by: Wang, Zhongdao, et al.
Published: (2025)
by: Wang, Zhongdao, et al.
Published: (2025)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025)
by: Ravichander, Abhilasha, et al.
Published: (2025)
ARGS: Alignment as Reward-Guided Search
by: Khanov, Maxim, et al.
Published: (2024)
by: Khanov, Maxim, et al.
Published: (2024)
Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them
by: Bui, Anh, et al.
Published: (2025)
by: Bui, Anh, et al.
Published: (2025)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
by: Patel, Fagun, et al.
Published: (2025)
by: Patel, Fagun, et al.
Published: (2025)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024)
by: Hardy, Amelia, et al.
Published: (2024)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAM
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
Neural Nonmyopic Bayesian Optimization in Dynamic Cost Settings
by: Truong, Sang T., et al.
Published: (2026)
by: Truong, Sang T., et al.
Published: (2026)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Fantastic Flips and Where to Find Them: A General Framework for Parameterized Local Search on Partitioning Problems
by: Grüttemeier, Niels, et al.
Published: (2025)
by: Grüttemeier, Niels, et al.
Published: (2025)
Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics
by: Liu, Zhu, et al.
Published: (2024)
by: Liu, Zhu, et al.
Published: (2024)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023)
by: Lamparth, Max, et al.
Published: (2023)
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Prediction of Item Difficulty for Reading Comprehension Items by Creation of Annotated Item Repository
by: Kapoor, Radhika, et al.
Published: (2025)
by: Kapoor, Radhika, et al.
Published: (2025)
Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model
by: Roth, Karsten, et al.
Published: (2023)
by: Roth, Karsten, et al.
Published: (2023)
Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
by: Ramtoula, Benjamin, et al.
Published: (2025)
by: Ramtoula, Benjamin, et al.
Published: (2025)
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
by: Chen, Edward, et al.
Published: (2025)
by: Chen, Edward, et al.
Published: (2025)
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
by: Hassanpour, Negar, et al.
Published: (2025)
by: Hassanpour, Negar, et al.
Published: (2025)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency
by: Zhang, Yibo Jacky, et al.
Published: (2026)
by: Zhang, Yibo Jacky, et al.
Published: (2026)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
Fantastic Copyrighted Beasts and How (Not) to Generate Them
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
Qudit Designs and Where to Find Them
by: Anand, Namit, et al.
Published: (2026)
by: Anand, Namit, et al.
Published: (2026)
Facts and People: Where to Find Them.
by: Hirigoyen, Maria
Published: (1971)
by: Hirigoyen, Maria
Published: (1971)
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
by: Semnani, Sina J., et al.
Published: (2025)
by: Semnani, Sina J., et al.
Published: (2025)
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Similar Items
-
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
by: Hardy, Michael, et al.
Published: (2026) -
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
by: Haupt, Andreas, et al.
Published: (2026) -
Your Classifier Can Be Secretly a Likelihood-Based OOD Detector
by: Burapacheep, Jirayu, et al.
Published: (2024) -
Fantastic Pretraining Optimizers and Where to Find Them
by: Wen, Kaiyue, et al.
Published: (2025) -
Fantastic Biases (What are They) and Where to Find Them
by: Barriere, Valentin
Published: (2024)