Themisto: Jupyter-Based Runtime Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910912993034240 |
|---|---|
| author | Grotov, Konstantin Titov, Sergey |
| author_facet | Grotov, Konstantin Titov, Sergey |
| contents | In this work, we present a benchmark that consists of Jupyter notebooks development trajectories and allows measuring how large language models (LLMs) can leverage runtime information for predicting code output and code generation. We demonstrate that the current generation of LLMs performs poorly on these tasks and argue that there exists a significantly understudied domain in the development of code-based models, which involves incorporating the runtime context. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_12365 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Themisto: Jupyter-Based Runtime Benchmark Grotov, Konstantin Titov, Sergey Software Engineering Artificial Intelligence Machine Learning In this work, we present a benchmark that consists of Jupyter notebooks development trajectories and allows measuring how large language models (LLMs) can leverage runtime information for predicting code output and code generation. We demonstrate that the current generation of LLMs performs poorly on these tasks and argue that there exists a significantly understudied domain in the development of code-based models, which involves incorporating the runtime context. |
| title | Themisto: Jupyter-Based Runtime Benchmark |
| topic | Software Engineering Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2504.12365 |