The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Uzunoglu, Arda, Li, Tianjian, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Jailbreak Distillation: Renewable Safety Benchmarking
par: Zhang, Jingyu, et autres
Publié: (2025)
par: Zhang, Jingyu, et autres
Publié: (2025)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
par: Uzunoglu, Arda, et autres
Publié: (2026)
par: Uzunoglu, Arda, et autres
Publié: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
par: Jing, Huihao, et autres
Publié: (2026)
par: Jing, Huihao, et autres
Publié: (2026)
Is Your Benchmark (Still) Useful? Dynamic Benchmarking for Code Language Models
par: Guan, Batu, et autres
Publié: (2025)
par: Guan, Batu, et autres
Publié: (2025)
Privacy Policy Analysis through Prompt Engineering for LLMs
par: Goknil, Arda, et autres
Publié: (2024)
par: Goknil, Arda, et autres
Publié: (2024)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
par: Xie, Yiqing, et autres
Publié: (2024)
par: Xie, Yiqing, et autres
Publié: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
Conjecture and Inquiry: Quantifying Software Performance Requirements via Interactive Retrieval-Augmented Preference Elicitation
par: Wang, Shihai, et autres
Publié: (2026)
par: Wang, Shihai, et autres
Publié: (2026)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
par: Du, Junjia, et autres
Publié: (2025)
par: Du, Junjia, et autres
Publié: (2025)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
par: Zhao, Songwen, et autres
Publié: (2025)
par: Zhao, Songwen, et autres
Publié: (2025)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
par: Li, Wei, et autres
Publié: (2025)
par: Li, Wei, et autres
Publié: (2025)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
par: Wang, Sizhe, et autres
Publié: (2025)
par: Wang, Sizhe, et autres
Publié: (2025)
Benchmarking Failures in Tool-Augmented Language Models
par: Treviño, Eduardo, et autres
Publié: (2025)
par: Treviño, Eduardo, et autres
Publié: (2025)
Benchmarks as Microscopes: A Call for Model Metrology
par: Saxon, Michael, et autres
Publié: (2024)
par: Saxon, Michael, et autres
Publié: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
par: Liu, Kaiyuan, et autres
Publié: (2025)
par: Liu, Kaiyuan, et autres
Publié: (2025)
A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
par: Guo, Jiale, et autres
Publié: (2025)
par: Guo, Jiale, et autres
Publié: (2025)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
par: Ou, Jiefu, et autres
Publié: (2024)
par: Ou, Jiefu, et autres
Publié: (2024)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
par: Huang, Dong, et autres
Publié: (2024)
par: Huang, Dong, et autres
Publié: (2024)
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
par: Zhu, Wang Bill, et autres
Publié: (2026)
par: Zhu, Wang Bill, et autres
Publié: (2026)
CommitBench: A Benchmark for Commit Message Generation
par: Schall, Maximilian, et autres
Publié: (2024)
par: Schall, Maximilian, et autres
Publié: (2024)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
par: Liu, Steven, et autres
Publié: (2026)
par: Liu, Steven, et autres
Publié: (2026)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
par: Huang, Yue, et autres
Publié: (2023)
par: Huang, Yue, et autres
Publié: (2023)
DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale
par: Zhang, Linghao, et autres
Publié: (2025)
par: Zhang, Linghao, et autres
Publié: (2025)
Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
par: Zhang, Daoan, et autres
Publié: (2026)
par: Zhang, Daoan, et autres
Publié: (2026)
Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence
par: Song, Linxin, et autres
Publié: (2026)
par: Song, Linxin, et autres
Publié: (2026)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
par: Liu, Zeyu Leo, et autres
Publié: (2024)
par: Liu, Zeyu Leo, et autres
Publié: (2024)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
par: Huang, Dong, et autres
Publié: (2025)
par: Huang, Dong, et autres
Publié: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
par: Chen, Zaoyu, et autres
Publié: (2026)
par: Chen, Zaoyu, et autres
Publié: (2026)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
par: Du, Mingzhe, et autres
Publié: (2024)
par: Du, Mingzhe, et autres
Publié: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
par: Peng, Qiwei, et autres
Publié: (2024)
par: Peng, Qiwei, et autres
Publié: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
par: Fu, Lingyue, et autres
Publié: (2025)
par: Fu, Lingyue, et autres
Publié: (2025)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
par: Chou, Jason, et autres
Publié: (2025)
par: Chou, Jason, et autres
Publié: (2025)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
par: Wu, Jiarong, et autres
Publié: (2025)
par: Wu, Jiarong, et autres
Publié: (2025)
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
par: Ding, Xianzhong, et autres
Publié: (2026)
par: Ding, Xianzhong, et autres
Publié: (2026)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
par: Li, Zhuohao, et autres
Publié: (2025)
par: Li, Zhuohao, et autres
Publié: (2025)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
par: Zheng, Dewu, et autres
Publié: (2024)
par: Zheng, Dewu, et autres
Publié: (2024)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
par: Chen, Yeheng, et autres
Publié: (2026)
par: Chen, Yeheng, et autres
Publié: (2026)
AlloyInEcore: Embedding of First-Order Relational Logic into Meta-Object Facility for Automated Model Reasoning
par: Erata, Ferhat, et autres
Publié: (2024)
par: Erata, Ferhat, et autres
Publié: (2024)
A Tool for Automated Reasoning About Traces Based on Configurable Formal Semantics
par: Erata, Ferhat, et autres
Publié: (2024)
par: Erata, Ferhat, et autres
Publié: (2024)
Documents similaires
-
Jailbreak Distillation: Renewable Safety Benchmarking
par: Zhang, Jingyu, et autres
Publié: (2025) -
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
par: Uzunoglu, Arda, et autres
Publié: (2026) -
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
par: Jing, Huihao, et autres
Publié: (2026) -
Is Your Benchmark (Still) Useful? Dynamic Benchmarking for Code Language Models
par: Guan, Batu, et autres
Publié: (2025) -
Privacy Policy Analysis through Prompt Engineering for LLMs
par: Goknil, Arda, et autres
Publié: (2024)