Gistify! Codebase-Level Understanding via Runtime Execution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914124613550080 |
|---|---|
| author | Lee, Hyunji Kim, Minseon Singh, Chinmay Pereira, Matheus Sonwane, Atharv White, Isadora Stengel-Eskin, Elias Bansal, Mohit Shi, Zhengyan Sordoni, Alessandro Côté, Marc-Alexandre Yuan, Xingdi Caccia, Lucas |
| author_facet | Lee, Hyunji Kim, Minseon Singh, Chinmay Pereira, Matheus Sonwane, Atharv White, Isadora Stengel-Eskin, Elias Bansal, Mohit Shi, Zhengyan Sordoni, Alessandro Côté, Marc-Alexandre Yuan, Xingdi Caccia, Lucas |
| contents | As coding agents are increasingly deployed in large codebases, the need to automatically design challenging, codebase-level evaluation is central. We propose Gistify, a task where a coding LLM must create a single, minimal, self-contained file that can reproduce a specific functionality of a codebase. The coding LLM is given full access to a codebase along with a specific entrypoint (e.g., a python command), and the generated file must replicate the output of the same command ran under the full codebase, while containing only the essential components necessary to execute the provided command. Success on Gistify requires both structural understanding of the codebase, accurate modeling of its execution flow as well as the ability to produce potentially large code patches. Our findings show that current state-of-the-art models struggle to reliably solve Gistify tasks, especially ones with long executions traces. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_26790 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Gistify! Codebase-Level Understanding via Runtime Execution Lee, Hyunji Kim, Minseon Singh, Chinmay Pereira, Matheus Sonwane, Atharv White, Isadora Stengel-Eskin, Elias Bansal, Mohit Shi, Zhengyan Sordoni, Alessandro Côté, Marc-Alexandre Yuan, Xingdi Caccia, Lucas Computation and Language Artificial Intelligence As coding agents are increasingly deployed in large codebases, the need to automatically design challenging, codebase-level evaluation is central. We propose Gistify, a task where a coding LLM must create a single, minimal, self-contained file that can reproduce a specific functionality of a codebase. The coding LLM is given full access to a codebase along with a specific entrypoint (e.g., a python command), and the generated file must replicate the output of the same command ran under the full codebase, while containing only the essential components necessary to execute the provided command. Success on Gistify requires both structural understanding of the codebase, accurate modeling of its execution flow as well as the ability to produce potentially large code patches. Our findings show that current state-of-the-art models struggle to reliably solve Gistify tasks, especially ones with long executions traces. |
| title | Gistify! Codebase-Level Understanding via Runtime Execution |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.26790 |