GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lindenbauer, Tobias, Bogomolov, Egor, Zharov, Yaroslav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912400186277888
author Lindenbauer, Tobias
Bogomolov, Egor
Zharov, Yaroslav
author_facet Lindenbauer, Tobias
Bogomolov, Egor
Zharov, Yaroslav
contents Benchmarks for Software Engineering (SE) AI agents, most notably SWE-bench, have catalyzed progress in programming capabilities of AI agents. However, they overlook critical developer workflows such as Version Control System (VCS) operations. To address this issue, we present GitGoodBench, a novel benchmark for evaluating AI agent performance on VCS tasks. GitGoodBench covers three core Git scenarios extracted from permissive open-source Python, Java, and Kotlin repositories. Our benchmark provides three datasets: a comprehensive evaluation suite (900 samples), a rapid prototyping version (120 samples), and a training corpus (17,469 samples). We establish baseline performance on the prototyping version of our benchmark using GPT-4o equipped with custom tools, achieving a 21.11% solve rate overall. We expect GitGoodBench to serve as a crucial stepping stone toward truly comprehensive SE agents that go beyond mere programming.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22583
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
Lindenbauer, Tobias
Bogomolov, Egor
Zharov, Yaroslav
Software Engineering
Artificial Intelligence
Benchmarks for Software Engineering (SE) AI agents, most notably SWE-bench, have catalyzed progress in programming capabilities of AI agents. However, they overlook critical developer workflows such as Version Control System (VCS) operations. To address this issue, we present GitGoodBench, a novel benchmark for evaluating AI agent performance on VCS tasks. GitGoodBench covers three core Git scenarios extracted from permissive open-source Python, Java, and Kotlin repositories. Our benchmark provides three datasets: a comprehensive evaluation suite (900 samples), a rapid prototyping version (120 samples), and a training corpus (17,469 samples). We establish baseline performance on the prototyping version of our benchmark using GPT-4o equipped with custom tools, achieving a 21.11% solve rate overall. We expect GitGoodBench to serve as a crucial stepping stone toward truly comprehensive SE agents that go beyond mere programming.
title GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2505.22583