CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Zhenru, Yao, Yiqun, Yuan, Yang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916145752178688
author Lin, Zhenru
Yao, Yiqun
Yuan, Yang
author_facet Lin, Zhenru
Yao, Yiqun
Yuan, Yang
contents Large language models (LLMs) such as ChatGPT are increasingly proficient in understanding and generating a mixture of code and text. Evaluation based on such $\textit{mixture}$ can lead to a more comprehensive understanding of the models' abilities in solving coding problems. However, in this context, current evaluation methods are either limited in task coverage or lack standardization. To address this issue, we propose using category theory as a framework for evaluation. Specifically, morphisms within a code category can represent code debugging and transformation, functors between two categories represent code translation, and functors between a code category and a natural language category represent code generation, explanation, and reproduction. We present an automatic evaluation framework called $\textbf{CatCode}$ ($\textbf{Cat}$egory $\textbf{Code}$) that can comprehensively assess the coding abilities of LLMs, including ChatGPT, Text-Davinci, and CodeGeeX.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01784
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text
Lin, Zhenru
Yao, Yiqun
Yuan, Yang
Artificial Intelligence
Programming Languages
Large language models (LLMs) such as ChatGPT are increasingly proficient in understanding and generating a mixture of code and text. Evaluation based on such $\textit{mixture}$ can lead to a more comprehensive understanding of the models' abilities in solving coding problems. However, in this context, current evaluation methods are either limited in task coverage or lack standardization. To address this issue, we propose using category theory as a framework for evaluation. Specifically, morphisms within a code category can represent code debugging and transformation, functors between two categories represent code translation, and functors between a code category and a natural language category represent code generation, explanation, and reproduction. We present an automatic evaluation framework called $\textbf{CatCode}$ ($\textbf{Cat}$egory $\textbf{Code}$) that can comprehensively assess the coding abilities of LLMs, including ChatGPT, Text-Davinci, and CodeGeeX.
title CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text
topic Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2403.01784