When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akhtar, Mubashara, Reuel, Anka, Soni, Prajna, Ahuja, Sanchit, Ammanamanchi, Pawan Sasanka, Rawal, Ruchit, Zouhar, Vilém, Yadav, Srishti, Whitehouse, Chenxi, Ki, Dayeon, Mickel, Jennifer, Choshen, Leshem, Šuppa, Marek, Batzner, Jan, Chim, Jenny, Sania, Jeba, Long, Yanan, Rahmani, Hossein A., Knight, Christina, Nan, Yiyang, Raj, Jyoutir, Fan, Yu, Singh, Shubham, Sahoo, Subramanyam, Habba, Eliya, Gohar, Usman, Pawar, Siddhesh, Scholz, Robert, Subramonian, Arjun, Ni, Jingwei, Kochenderfer, Mykel, Koyejo, Sanmi, Sachan, Mrinmaya, Biderman, Stella, Talat, Zeerak, Ghosh, Avijit, Solaiman, Irene |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
von: Reuel, Anka, et al.
Veröffentlicht: (2025)
von: Reuel, Anka, et al.
Veröffentlicht: (2025)
Beyond Benchmarks: On The False Promise of AI Regulation
von: Stanovsky, Gabriel, et al.
Veröffentlicht: (2025)
von: Stanovsky, Gabriel, et al.
Veröffentlicht: (2025)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
von: Zouhar, Vilém
Veröffentlicht: (2024)
von: Zouhar, Vilém
Veröffentlicht: (2024)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
BlasBench: An Open Benchmark for Irish Speech Recognition
von: Raj, Jyoutir, et al.
Veröffentlicht: (2026)
von: Raj, Jyoutir, et al.
Veröffentlicht: (2026)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
von: Lamparth, Max, et al.
Veröffentlicht: (2023)
von: Lamparth, Max, et al.
Veröffentlicht: (2023)
Fairness in Reinforcement Learning: A Survey
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Pearmut: Human Evaluation of Translation Made Trivial
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
von: Haupt, Andreas, et al.
Veröffentlicht: (2026)
von: Haupt, Andreas, et al.
Veröffentlicht: (2026)
Generative AI Needs Adaptive Governance
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
JSON Whisperer: Efficient JSON Editing with LLMs
von: Duanis, Sarel, et al.
Veröffentlicht: (2025)
von: Duanis, Sarel, et al.
Veröffentlicht: (2025)
Understanding "Democratization" in NLP and ML Research
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
Distributional Properties of Subword Regularization
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
AI-Assisted Human Evaluation of Machine Translation
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
Multimodal Shannon Game with Images
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
Time-Dependent Queuing Model for Traffic Congestion Using Mt/D/1/K: Simulation and Policy Insights
von: Raj, Jyoutir
Veröffentlicht: (2025)
von: Raj, Jyoutir
Veröffentlicht: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
von: Chowdhury, Sankalan Pal, et al.
Veröffentlicht: (2024)
von: Chowdhury, Sankalan Pal, et al.
Veröffentlicht: (2024)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
von: Lior, Gili, et al.
Veröffentlicht: (2025)
von: Lior, Gili, et al.
Veröffentlicht: (2025)
Audit Cards: Contextualizing AI Evaluations
von: Staufer, Leon, et al.
Veröffentlicht: (2025)
von: Staufer, Leon, et al.
Veröffentlicht: (2025)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
von: Rooein, Donya, et al.
Veröffentlicht: (2025)
von: Rooein, Donya, et al.
Veröffentlicht: (2025)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
Evaluating Optimal Reference Translations
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
A Bayesian Optimization Approach to Machine Translation Reranking
von: Cheng, Julius, et al.
Veröffentlicht: (2024)
von: Cheng, Julius, et al.
Veröffentlicht: (2024)
How to Engage Your Readers? Generating Guiding Questions to Promote Active Reading
von: Cui, Peng, et al.
Veröffentlicht: (2024)
von: Cui, Peng, et al.
Veröffentlicht: (2024)
Two Counterexamples to Tokenization and the Noiseless Channel
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
von: Hardy, Michael, et al.
Veröffentlicht: (2026)
von: Hardy, Michael, et al.
Veröffentlicht: (2026)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
von: Hardy, Amelia, et al.
Veröffentlicht: (2024)
von: Hardy, Amelia, et al.
Veröffentlicht: (2024)
Fantastic Bugs and Where to Find Them in AI Benchmarks
von: Truong, Sang, et al.
Veröffentlicht: (2025)
von: Truong, Sang, et al.
Veröffentlicht: (2025)
CafGa: Customizing Feature Attributions to Explain Language Models
von: Boyle, Alan, et al.
Veröffentlicht: (2025)
von: Boyle, Alan, et al.
Veröffentlicht: (2025)
Agent Benchmarks Fail Public Sector Requirements
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2026)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2026)
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
von: Levy, Shahar, et al.
Veröffentlicht: (2026)
von: Levy, Shahar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
von: Reuel, Anka, et al.
Veröffentlicht: (2025) -
Beyond Benchmarks: On The False Promise of AI Regulation
von: Stanovsky, Gabriel, et al.
Veröffentlicht: (2025) -
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
von: Zouhar, Vilém
Veröffentlicht: (2024) -
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026) -
BlasBench: An Open Benchmark for Irish Speech Recognition
von: Raj, Jyoutir, et al.
Veröffentlicht: (2026)