When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Shenyu, Dong, Ximing, Liu, Xiaoshuang, Oliva, Gustavo, Yong, Chong Chun, Lin, Dayi, Chen, Boyuan, Wang, Shaowei, Hassan, Ahmed E. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024)
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024)
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024)
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024)
Does Using Bazel Help Speed Up Continuous Integration Builds?
di: Zheng, Shenyu, et al.
Pubblicazione: (2024)
di: Zheng, Shenyu, et al.
Pubblicazione: (2024)
Code Generation with Small Language Models: A Codeforces-Based Study
di: Souza, Débora, et al.
Pubblicazione: (2025)
di: Souza, Débora, et al.
Pubblicazione: (2025)
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
di: Chen, Zhilong, et al.
Pubblicazione: (2025)
di: Chen, Zhilong, et al.
Pubblicazione: (2025)
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
di: Rajbahadur, Gopi Krishnan, et al.
Pubblicazione: (2024)
di: Rajbahadur, Gopi Krishnan, et al.
Pubblicazione: (2024)
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
di: Wang, Chaozheng, et al.
Pubblicazione: (2024)
di: Wang, Chaozheng, et al.
Pubblicazione: (2024)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
di: Oliva, Gustavo A., et al.
Pubblicazione: (2025)
di: Oliva, Gustavo A., et al.
Pubblicazione: (2025)
Predicting long time contributors with knowledge units of programming languages: an empirical study
di: Ahasanuzzaman, Md, et al.
Pubblicazione: (2024)
di: Ahasanuzzaman, Md, et al.
Pubblicazione: (2024)
SynConfRoute: Syntax-Aware Routing for Efficient Code Completion with Small CodeLLMs
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2026)
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2026)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
di: Fan, Zhiyu, et al.
Pubblicazione: (2025)
di: Fan, Zhiyu, et al.
Pubblicazione: (2025)
Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
di: Cogo, Filipe R., et al.
Pubblicazione: (2025)
di: Cogo, Filipe R., et al.
Pubblicazione: (2025)
Engineering AI Judge Systems
di: Lin, Jiahuei, et al.
Pubblicazione: (2024)
di: Lin, Jiahuei, et al.
Pubblicazione: (2024)
SLA-Awareness for AI-assisted coding
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2025)
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2025)
Context-Aware CodeLLM Eviction for AI-assisted Coding
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2025)
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2025)
Predicting post-release defects with knowledge units (KUs) of programming languages: an empirical study
di: Ahasanuzzaman, Md, et al.
Pubblicazione: (2024)
di: Ahasanuzzaman, Md, et al.
Pubblicazione: (2024)
When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation
di: Ashik, Ahmed Nusayer, et al.
Pubblicazione: (2026)
di: Ashik, Ahmed Nusayer, et al.
Pubblicazione: (2026)
Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study
di: Ahasanuzzaman, Md, et al.
Pubblicazione: (2026)
di: Ahasanuzzaman, Md, et al.
Pubblicazione: (2026)
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024)
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
di: Feng, Jia, et al.
Pubblicazione: (2024)
di: Feng, Jia, et al.
Pubblicazione: (2024)
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement
di: Gallaba, Keheliya, et al.
Pubblicazione: (2025)
di: Gallaba, Keheliya, et al.
Pubblicazione: (2025)
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
di: Rombaut, Benjamin, et al.
Pubblicazione: (2024)
di: Rombaut, Benjamin, et al.
Pubblicazione: (2024)
Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees
di: Chang, Shi, et al.
Pubblicazione: (2025)
di: Chang, Shi, et al.
Pubblicazione: (2025)
Towards Reliable Generation of Executable Workflows by Foundation Models
di: Masoumzadeh, Sogol, et al.
Pubblicazione: (2025)
di: Masoumzadeh, Sogol, et al.
Pubblicazione: (2025)
When Model Editing Meets Service Evolution: A Knowledge-Update Perspective for Service Recommendation
di: Fan, Guodong, et al.
Pubblicazione: (2026)
di: Fan, Guodong, et al.
Pubblicazione: (2026)
A Large-Scale Exploratory Study on the Proxy Pattern in Ethereum
di: Ebrahimi, Amir M., et al.
Pubblicazione: (2025)
di: Ebrahimi, Amir M., et al.
Pubblicazione: (2025)
Data Quality Antipatterns for Software Analytics
di: Bhatia, Aaditya, et al.
Pubblicazione: (2024)
di: Bhatia, Aaditya, et al.
Pubblicazione: (2024)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
di: Hassan, Ahmed E., et al.
Pubblicazione: (2025)
di: Hassan, Ahmed E., et al.
Pubblicazione: (2025)
Open Source, Hidden Costs: A Systematic Literature Review on OSS License Management
di: Li, Boyuan, et al.
Pubblicazione: (2025)
di: Li, Boyuan, et al.
Pubblicazione: (2025)
Codehacks: A Dataset of Adversarial Tests for Competitive Programming Problems Obtained from Codeforces
di: Hort, Max, et al.
Pubblicazione: (2025)
di: Hort, Max, et al.
Pubblicazione: (2025)
Can LLMs be Effective Code Contributors? A Study on Open-source Projects
di: Chong, Chun Jie, et al.
Pubblicazione: (2026)
di: Chong, Chun Jie, et al.
Pubblicazione: (2026)
MCeT: Behavioral Model Correctness Evaluation using Large Language Models
di: Ahmed, Khaled, et al.
Pubblicazione: (2025)
di: Ahmed, Khaled, et al.
Pubblicazione: (2025)
TEASMA: A Practical Methodology for Test Adequacy Assessment of Deep Neural Networks
di: Abbasishahkoo, Amin, et al.
Pubblicazione: (2023)
di: Abbasishahkoo, Amin, et al.
Pubblicazione: (2023)
Evaluating the Effectiveness and Efficiency of Demonstration Retrievers in RAG for Coding Tasks
di: He, Pengfei, et al.
Pubblicazione: (2024)
di: He, Pengfei, et al.
Pubblicazione: (2024)
Output Format Biases in the Evaluation of Large Language Models for Code Translation
di: Macedo, Marcos, et al.
Pubblicazione: (2024)
di: Macedo, Marcos, et al.
Pubblicazione: (2024)
UPC Sentinel: An Accurate Approach for Detecting Upgradeability Proxy Contracts in Ethereum
di: Ebrahimi, Amir M., et al.
Pubblicazione: (2024)
di: Ebrahimi, Amir M., et al.
Pubblicazione: (2024)
SimClone: Detecting Tabular Data Clones using Value Similarity
di: Yang, Xu, et al.
Pubblicazione: (2024)
di: Yang, Xu, et al.
Pubblicazione: (2024)
Can We Recycle Our Old Models? An Empirical Evaluation of Model Selection Mechanisms for AIOps Solutions
di: Lyu, Yingzhe, et al.
Pubblicazione: (2025)
di: Lyu, Yingzhe, et al.
Pubblicazione: (2025)
A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
di: Khan, Taufiqul Islam, et al.
Pubblicazione: (2026)
di: Khan, Taufiqul Islam, et al.
Pubblicazione: (2026)
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
di: Dai, Dekun, et al.
Pubblicazione: (2025)
di: Dai, Dekun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024) -
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
di: Hassan, Ahmed E., et al.
Pubblicazione: (2024) -
Does Using Bazel Help Speed Up Continuous Integration Builds?
di: Zheng, Shenyu, et al.
Pubblicazione: (2024) -
Code Generation with Small Language Models: A Codeforces-Based Study
di: Souza, Débora, et al.
Pubblicazione: (2025) -
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
di: Chen, Zhilong, et al.
Pubblicazione: (2025)