KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lai, Eugenie, Vitagliano, Gerardo, Zhang, Ziyu, Chabra, Om, Sudhir, Sivaprasad, Zeng, Anna, Zabreyko, Anton A., Li, Chenning, Kossmann, Ferdi, Ding, Jialin, Chen, Jun, Markakis, Markos, Russo, Matthew, Wang, Weiyang, Wu, Ziniu, Cafarella, Michael J., Cao, Lei, Madden, Samuel, Kraska, Tim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917316335239168
author Lai, Eugenie
Vitagliano, Gerardo
Zhang, Ziyu
Chabra, Om
Sudhir, Sivaprasad
Zeng, Anna
Zabreyko, Anton A.
Li, Chenning
Kossmann, Ferdi
Ding, Jialin
Chen, Jun
Markakis, Markos
Russo, Matthew
Wang, Weiyang
Wu, Ziniu
Cafarella, Michael J.
Cao, Lei
Madden, Samuel
Kraska, Tim
author_facet Lai, Eugenie
Vitagliano, Gerardo
Zhang, Ziyu
Chabra, Om
Sudhir, Sivaprasad
Zeng, Anna
Zabreyko, Anton A.
Li, Chenning
Kossmann, Ferdi
Ding, Jialin
Chen, Jun
Markakis, Markos
Russo, Matthew
Wang, Weiyang
Wu, Ziniu
Cafarella, Michael J.
Cao, Lei
Madden, Samuel
Kraska, Tim
contents Discovering insights from a real-world data lake potentially containing unclean, semi-structured, and unstructured data requires a variety of data processing tasks, ranging from extraction and cleaning to integration, analysis, and modeling. This process often also demands domain knowledge and project-specific insight. While AI models have shown remarkable results in reasoning and code generation, their abilities to design and execute complex pipelines that solve these data-lake-to-insight challenges remain unclear. We introduce KramaBench which consists of 104 manually curated and solved challenges spanning 1700 files, 24 data sources, and 6 domains. KramaBench focuses on testing the end-to-end capabilities of AI systems to solve challenges which require automated orchestration of different data tasks. KramaBench also features a comprehensive evaluation framework assessing the pipeline design and individual data task implementation abilities of AI systems. We evaluate 8 LLMs using our single-agent reference framework DS-Guru, alongside both open- and closed-source single- and multi-agent systems, and find that while current agentic systems may handle isolated data-science tasks and generate plausible draft pipelines, they struggle with producing working end-to-end pipelines. On KramaBench, the best system reaches only 55% end-to-end accuracy in the full data-lake setting. Even with perfect retrieval, the accuracy tops out at 62%. Leading LLMs can identify up to 42% of important data tasks but can only fully implement 20% of individual data tasks. Our code, reference framework, and data are available at https://github.com/mitdbg/KramaBench.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes
Lai, Eugenie
Vitagliano, Gerardo
Zhang, Ziyu
Chabra, Om
Sudhir, Sivaprasad
Zeng, Anna
Zabreyko, Anton A.
Li, Chenning
Kossmann, Ferdi
Ding, Jialin
Chen, Jun
Markakis, Markos
Russo, Matthew
Wang, Weiyang
Wu, Ziniu
Cafarella, Michael J.
Cao, Lei
Madden, Samuel
Kraska, Tim
Databases
Artificial Intelligence
Multiagent Systems
Discovering insights from a real-world data lake potentially containing unclean, semi-structured, and unstructured data requires a variety of data processing tasks, ranging from extraction and cleaning to integration, analysis, and modeling. This process often also demands domain knowledge and project-specific insight. While AI models have shown remarkable results in reasoning and code generation, their abilities to design and execute complex pipelines that solve these data-lake-to-insight challenges remain unclear. We introduce KramaBench which consists of 104 manually curated and solved challenges spanning 1700 files, 24 data sources, and 6 domains. KramaBench focuses on testing the end-to-end capabilities of AI systems to solve challenges which require automated orchestration of different data tasks. KramaBench also features a comprehensive evaluation framework assessing the pipeline design and individual data task implementation abilities of AI systems. We evaluate 8 LLMs using our single-agent reference framework DS-Guru, alongside both open- and closed-source single- and multi-agent systems, and find that while current agentic systems may handle isolated data-science tasks and generate plausible draft pipelines, they struggle with producing working end-to-end pipelines. On KramaBench, the best system reaches only 55% end-to-end accuracy in the full data-lake setting. Even with perfect retrieval, the accuracy tops out at 62%. Leading LLMs can identify up to 42% of important data tasks but can only fully implement 20% of individual data tasks. Our code, reference framework, and data are available at https://github.com/mitdbg/KramaBench.
title KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes
topic Databases
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2506.06541