CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Manh, Dung Nguyen, Chau, Thang Phan, Hai, Nam Le, Doan, Thong T., Nguyen, Nam V., Pham, Quang, Bui, Nghi D. Q.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915234362425344
author Manh, Dung Nguyen
Chau, Thang Phan
Hai, Nam Le
Doan, Thong T.
Nguyen, Nam V.
Pham, Quang
Bui, Nghi D. Q.
author_facet Manh, Dung Nguyen
Chau, Thang Phan
Hai, Nam Le
Doan, Thong T.
Nguyen, Nam V.
Pham, Quang
Bui, Nghi D. Q.
contents Recent advances in Code Large Language Models (CodeLLMs) have primarily focused on open-ended code generation, often overlooking the crucial aspect of code understanding and reasoning. To bridge this gap, we introduce CodeMMLU, a comprehensive multiple-choice benchmark designed to evaluate the depth of software and code comprehension in LLMs. CodeMMLU includes nearly 20,000 questions spanning diverse domains, including code analysis, defect detection, and software engineering principles across multiple programming languages. Unlike traditional benchmarks that emphasize code generation, CodeMMLU assesses a model's ability to reason about programs across a wide-range of tasks such as code repair, execution reasoning, and fill-in-the-blank challenges. Our extensive evaluation reveals that even state-of-the-art models struggle with CodeMMLU, highlighting significant gaps in comprehension beyond generation. By emphasizing the essential connection between code understanding and effective AI-assisted development, CodeMMLU provides a critical resource for advancing more reliable and capable coding assistants.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01999
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
Manh, Dung Nguyen
Chau, Thang Phan
Hai, Nam Le
Doan, Thong T.
Nguyen, Nam V.
Pham, Quang
Bui, Nghi D. Q.
Software Engineering
Recent advances in Code Large Language Models (CodeLLMs) have primarily focused on open-ended code generation, often overlooking the crucial aspect of code understanding and reasoning. To bridge this gap, we introduce CodeMMLU, a comprehensive multiple-choice benchmark designed to evaluate the depth of software and code comprehension in LLMs. CodeMMLU includes nearly 20,000 questions spanning diverse domains, including code analysis, defect detection, and software engineering principles across multiple programming languages. Unlike traditional benchmarks that emphasize code generation, CodeMMLU assesses a model's ability to reason about programs across a wide-range of tasks such as code repair, execution reasoning, and fill-in-the-blank challenges. Our extensive evaluation reveals that even state-of-the-art models struggle with CodeMMLU, highlighting significant gaps in comprehension beyond generation. By emphasizing the essential connection between code understanding and effective AI-assisted development, CodeMMLU provides a critical resource for advancing more reliable and capable coding assistants.
title CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
topic Software Engineering
url https://arxiv.org/abs/2410.01999