Gistify! Codebase-Level Understanding via Runtime Execution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Hyunji, Kim, Minseon, Singh, Chinmay, Pereira, Matheus, Sonwane, Atharv, White, Isadora, Stengel-Eskin, Elias, Bansal, Mohit, Shi, Zhengyan, Sordoni, Alessandro, Côté, Marc-Alexandre, Yuan, Xingdi, Caccia, Lucas
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914124613550080
author Lee, Hyunji
Kim, Minseon
Singh, Chinmay
Pereira, Matheus
Sonwane, Atharv
White, Isadora
Stengel-Eskin, Elias
Bansal, Mohit
Shi, Zhengyan
Sordoni, Alessandro
Côté, Marc-Alexandre
Yuan, Xingdi
Caccia, Lucas
author_facet Lee, Hyunji
Kim, Minseon
Singh, Chinmay
Pereira, Matheus
Sonwane, Atharv
White, Isadora
Stengel-Eskin, Elias
Bansal, Mohit
Shi, Zhengyan
Sordoni, Alessandro
Côté, Marc-Alexandre
Yuan, Xingdi
Caccia, Lucas
contents As coding agents are increasingly deployed in large codebases, the need to automatically design challenging, codebase-level evaluation is central. We propose Gistify, a task where a coding LLM must create a single, minimal, self-contained file that can reproduce a specific functionality of a codebase. The coding LLM is given full access to a codebase along with a specific entrypoint (e.g., a python command), and the generated file must replicate the output of the same command ran under the full codebase, while containing only the essential components necessary to execute the provided command. Success on Gistify requires both structural understanding of the codebase, accurate modeling of its execution flow as well as the ability to produce potentially large code patches. Our findings show that current state-of-the-art models struggle to reliably solve Gistify tasks, especially ones with long executions traces.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26790
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gistify! Codebase-Level Understanding via Runtime Execution
Lee, Hyunji
Kim, Minseon
Singh, Chinmay
Pereira, Matheus
Sonwane, Atharv
White, Isadora
Stengel-Eskin, Elias
Bansal, Mohit
Shi, Zhengyan
Sordoni, Alessandro
Côté, Marc-Alexandre
Yuan, Xingdi
Caccia, Lucas
Computation and Language
Artificial Intelligence
As coding agents are increasingly deployed in large codebases, the need to automatically design challenging, codebase-level evaluation is central. We propose Gistify, a task where a coding LLM must create a single, minimal, self-contained file that can reproduce a specific functionality of a codebase. The coding LLM is given full access to a codebase along with a specific entrypoint (e.g., a python command), and the generated file must replicate the output of the same command ran under the full codebase, while containing only the essential components necessary to execute the provided command. Success on Gistify requires both structural understanding of the codebase, accurate modeling of its execution flow as well as the ability to produce potentially large code patches. Our findings show that current state-of-the-art models struggle to reliably solve Gistify tasks, especially ones with long executions traces.
title Gistify! Codebase-Level Understanding via Runtime Execution
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.26790