GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
Fuente:
arXiv
Saved in:
| Main Authors: | Lindenbauer, Tobias, Bogomolov, Egor, Zharov, Yaroslav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026)
by: Slinko, Igor, et al.
Published: (2026)
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025)
by: Kovrigin, Alexander, et al.
Published: (2025)
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)
by: Eliseeva, Aleksandra, et al.
Published: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
On Problems of Implicit Context Compression for Software Engineering Agents
by: Gelvan, Kirill, et al.
Published: (2026)
by: Gelvan, Kirill, et al.
Published: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
Tool-Augmented LLMs as a Universal Interface for IDEs
by: Zharov, Yaroslav, et al.
Published: (2024)
by: Zharov, Yaroslav, et al.
Published: (2024)
GitHub Copilot: the perfect Code compLeeter?
by: Siroš, Ilja, et al.
Published: (2024)
by: Siroš, Ilja, et al.
Published: (2024)
AIDev: Studying AI Coding Agents on GitHub
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
by: Zan, Daoguang, et al.
Published: (2024)
by: Zan, Daoguang, et al.
Published: (2024)
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
by: Pandey, Ruchika, et al.
Published: (2024)
by: Pandey, Ruchika, et al.
Published: (2024)
Programming with AI: Evaluating ChatGPT, Gemini, AlphaCode, and GitHub Copilot for Programmers
by: Siam, Md Kamrul, et al.
Published: (2024)
by: Siam, Md Kamrul, et al.
Published: (2024)
Time Travel: LLM-Assisted Semantic Behavior Localization with Git Bisect
by: Wang, Yujing, et al.
Published: (2025)
by: Wang, Yujing, et al.
Published: (2025)
Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
From Knowledge to Noise: CTIM-Rover and the Pitfalls of Episodic Memory in Software Engineering Agents
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding
by: Nevo, Ziv, et al.
Published: (2025)
by: Nevo, Ziv, et al.
Published: (2025)
Using LLMs in Software Design: An Empirical Study of GitHub and A Practitioner Survey
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution
by: Tao, Wei, et al.
Published: (2024)
by: Tao, Wei, et al.
Published: (2024)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024)
by: Kovrigin, Alexander, et al.
Published: (2024)
RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
by: Wang, Huacan, et al.
Published: (2025)
by: Wang, Huacan, et al.
Published: (2025)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
Dynamic Retrieval-Augmented Generation
by: Shapkin, Anton, et al.
Published: (2023)
by: Shapkin, Anton, et al.
Published: (2023)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025)
by: Jewitt, James, et al.
Published: (2025)
Challenge on Optimization of Context Collection for Code Completion
by: Ustalov, Dmitry, et al.
Published: (2025)
by: Ustalov, Dmitry, et al.
Published: (2025)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
by: Glukhov, Evgeniy, et al.
Published: (2025)
by: Glukhov, Evgeniy, et al.
Published: (2025)
The Role of GitHub Copilot on Software Development: A Perspective on Productivity, Security, Best Practices and Future Directions
by: Nettur, Suresh Babu, et al.
Published: (2025)
by: Nettur, Suresh Babu, et al.
Published: (2025)
GitHub's Copilot Code Review: Can AI Spot Security Flaws Before You Commit?
by: Amro, Amena, et al.
Published: (2025)
by: Amro, Amena, et al.
Published: (2025)
The Impact of AI Tool on Engineering at ANZ Bank An Empirical Study on GitHub Copilot within Corporate Environment
by: Chatterjee, Sayan, et al.
Published: (2024)
by: Chatterjee, Sayan, et al.
Published: (2024)
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
by: Misra, Diganta, et al.
Published: (2025)
by: Misra, Diganta, et al.
Published: (2025)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
RAG4Tickets: AI-Powered Ticket Resolution via Retrieval-Augmented Generation on JIRA and GitHub Data
by: Baqar, Mohammad
Published: (2025)
by: Baqar, Mohammad
Published: (2025)
LLM-based Content Classification Approach for GitHub Repositories by the README Files
by: Mehmood, Malik Uzair, et al.
Published: (2025)
by: Mehmood, Malik Uzair, et al.
Published: (2025)
GitBug-Actions: Building Reproducible Bug-Fix Benchmarks with GitHub Actions
by: Saavedra, Nuno, et al.
Published: (2023)
by: Saavedra, Nuno, et al.
Published: (2023)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
by: Jimenez, Carlos E., et al.
Published: (2023)
by: Jimenez, Carlos E., et al.
Published: (2023)
GitEvo: Code Evolution Analysis for Git Repositories
by: Hora, Andre
Published: (2026)
by: Hora, Andre
Published: (2026)
AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Similar Items
-
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
by: Lindenbauer, Tobias, et al.
Published: (2025) -
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026) -
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025) -
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025) -
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)