Themisto: Jupyter-Based Runtime Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grotov, Konstantin, Titov, Sergey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910912993034240
author Grotov, Konstantin
Titov, Sergey
author_facet Grotov, Konstantin
Titov, Sergey
contents In this work, we present a benchmark that consists of Jupyter notebooks development trajectories and allows measuring how large language models (LLMs) can leverage runtime information for predicting code output and code generation. We demonstrate that the current generation of LLMs performs poorly on these tasks and argue that there exists a significantly understudied domain in the development of code-based models, which involves incorporating the runtime context.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Themisto: Jupyter-Based Runtime Benchmark
Grotov, Konstantin
Titov, Sergey
Software Engineering
Artificial Intelligence
Machine Learning
In this work, we present a benchmark that consists of Jupyter notebooks development trajectories and allows measuring how large language models (LLMs) can leverage runtime information for predicting code output and code generation. We demonstrate that the current generation of LLMs performs poorly on these tasks and argue that there exists a significantly understudied domain in the development of code-based models, which involves incorporating the runtime context.
title Themisto: Jupyter-Based Runtime Benchmark
topic Software Engineering
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.12365