CommitBench: A Benchmark for Commit Message Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Schall, Maximilian, Czinczoll, Tamara, de Melo, Gerard
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917608314372096
author Schall, Maximilian
Czinczoll, Tamara
de Melo, Gerard
author_facet Schall, Maximilian
Czinczoll, Tamara
de Melo, Gerard
contents Writing commit messages is a tedious daily task for many software developers, and often remains neglected. Automating this task has the potential to save time while ensuring that messages are informative. A high-quality dataset and an objective benchmark are vital preconditions for solid research and evaluation towards this goal. We show that existing datasets exhibit various problems, such as the quality of the commit selection, small sample sizes, duplicates, privacy issues, and missing licenses for redistribution. This can lead to unusable models and skewed evaluations, where inferior models achieve higher evaluation scores due to biases in the data. We compile a new large-scale dataset, CommitBench, adopting best practices for dataset creation. We sample commits from diverse projects with licenses that permit redistribution and apply our filtering and dataset enhancements to improve the quality of generated commit messages. We use CommitBench to compare existing models and show that other approaches are outperformed by a Transformer model pretrained on source code. We hope to accelerate future research by publishing the source code( https://github.com/Maxscha/commitbench ).
format Preprint
id arxiv_https___arxiv_org_abs_2403_05188
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CommitBench: A Benchmark for Commit Message Generation
Schall, Maximilian
Czinczoll, Tamara
de Melo, Gerard
Computation and Language
Software Engineering
Writing commit messages is a tedious daily task for many software developers, and often remains neglected. Automating this task has the potential to save time while ensuring that messages are informative. A high-quality dataset and an objective benchmark are vital preconditions for solid research and evaluation towards this goal. We show that existing datasets exhibit various problems, such as the quality of the commit selection, small sample sizes, duplicates, privacy issues, and missing licenses for redistribution. This can lead to unusable models and skewed evaluations, where inferior models achieve higher evaluation scores due to biases in the data. We compile a new large-scale dataset, CommitBench, adopting best practices for dataset creation. We sample commits from diverse projects with licenses that permit redistribution and apply our filtering and dataset enhancements to improve the quality of generated commit messages. We use CommitBench to compare existing models and show that other approaches are outperformed by a Transformer model pretrained on source code. We hope to accelerate future research by publishing the source code( https://github.com/Maxscha/commitbench ).
title CommitBench: A Benchmark for Commit Message Generation
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2403.05188