MATCH: Task-Driven Code Evaluation through Contrastive Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ghoummaid, Marah, Tchuiev, Vladimir, Glick, Ofek, Moshkovitz, Michal, Di Castro, Dotan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909872949297152
author Ghoummaid, Marah
Tchuiev, Vladimir
Glick, Ofek
Moshkovitz, Michal
Di Castro, Dotan
author_facet Ghoummaid, Marah
Tchuiev, Vladimir
Glick, Ofek
Moshkovitz, Michal
Di Castro, Dotan
contents AI-based code generation is increasingly prevalent, with GitHub Copilot estimated to generate 46% of the code on GitHub. Accurately evaluating how well generated code aligns with developer intent remains a critical challenge. Traditional evaluation methods, such as unit tests, are often unscalable and costly. Syntactic similarity metrics (e.g., BLEU, ROUGE) fail to capture code functionality, and metrics like CodeBERTScore require reference code, which is not always available. To address the gap in reference-free evaluation, with few alternatives such as ICE-Score, this paper introduces MATCH, a novel reference-free metric. MATCH uses Contrastive Learning to generate meaningful embeddings for code and natural language task descriptions, enabling similarity scoring that reflects how well generated code implements the task. We show that MATCH achieves stronger correlations with functional correctness and human preference than existing metrics across multiple programming languages.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23169
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MATCH: Task-Driven Code Evaluation through Contrastive Learning
Ghoummaid, Marah
Tchuiev, Vladimir
Glick, Ofek
Moshkovitz, Michal
Di Castro, Dotan
Computation and Language
Software Engineering
AI-based code generation is increasingly prevalent, with GitHub Copilot estimated to generate 46% of the code on GitHub. Accurately evaluating how well generated code aligns with developer intent remains a critical challenge. Traditional evaluation methods, such as unit tests, are often unscalable and costly. Syntactic similarity metrics (e.g., BLEU, ROUGE) fail to capture code functionality, and metrics like CodeBERTScore require reference code, which is not always available. To address the gap in reference-free evaluation, with few alternatives such as ICE-Score, this paper introduces MATCH, a novel reference-free metric. MATCH uses Contrastive Learning to generate meaningful embeddings for code and natural language task descriptions, enabling similarity scoring that reflects how well generated code implements the task. We show that MATCH achieves stronger correlations with functional correctness and human preference than existing metrics across multiple programming languages.
title MATCH: Task-Driven Code Evaluation through Contrastive Learning
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2510.23169