ConDABench: Interactive Evaluation of Language Models for Data Analysis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dutta, Avik, Gupta, Priyanshu, Hasanbeig, Hosein, Singh, Rahul Pratap, Nigam, Harshit, Gulwani, Sumit, Radhakrishna, Arjun, Soares, Gustavo, Tiwari, Ashish
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912651069620224
author Dutta, Avik
Gupta, Priyanshu
Hasanbeig, Hosein
Singh, Rahul Pratap
Nigam, Harshit
Gulwani, Sumit
Radhakrishna, Arjun
Soares, Gustavo
Tiwari, Ashish
author_facet Dutta, Avik
Gupta, Priyanshu
Hasanbeig, Hosein
Singh, Rahul Pratap
Nigam, Harshit
Gulwani, Sumit
Radhakrishna, Arjun
Soares, Gustavo
Tiwari, Ashish
contents Real-world data analysis tasks often come with under-specified goals and unclean data. User interaction is necessary to understand and disambiguate a user's intent, and hence, essential to solving these complex tasks. Existing benchmarks for evaluating LLMs on data analysis tasks do not capture these complexities or provide first-class support for interactivity. We introduce ConDABench, a framework for generating conversational data analysis (ConDA) benchmarks and evaluating external tools on the generated benchmarks. \bench consists of (a) a multi-agent workflow for generating realistic benchmarks from articles describing insights gained from public datasets, (b) 1,420 ConDA problems generated using this workflow, and (c) an evaluation harness that, for the first time, makes it possible to systematically evaluate conversational data analysis tools on the generated ConDA problems. Evaluation of state-of-the-art LLMs on the benchmarks reveals that while the new generation of models are better at solving more instances, they are not necessarily better at solving tasks that require sustained, long-form engagement. ConDABench is an avenue for model builders to measure progress towards truly collaborative models that can complete complex interactive tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13835
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConDABench: Interactive Evaluation of Language Models for Data Analysis
Dutta, Avik
Gupta, Priyanshu
Hasanbeig, Hosein
Singh, Rahul Pratap
Nigam, Harshit
Gulwani, Sumit
Radhakrishna, Arjun
Soares, Gustavo
Tiwari, Ashish
Computation and Language
Artificial Intelligence
Real-world data analysis tasks often come with under-specified goals and unclean data. User interaction is necessary to understand and disambiguate a user's intent, and hence, essential to solving these complex tasks. Existing benchmarks for evaluating LLMs on data analysis tasks do not capture these complexities or provide first-class support for interactivity. We introduce ConDABench, a framework for generating conversational data analysis (ConDA) benchmarks and evaluating external tools on the generated benchmarks. \bench consists of (a) a multi-agent workflow for generating realistic benchmarks from articles describing insights gained from public datasets, (b) 1,420 ConDA problems generated using this workflow, and (c) an evaluation harness that, for the first time, makes it possible to systematically evaluate conversational data analysis tools on the generated ConDA problems. Evaluation of state-of-the-art LLMs on the benchmarks reveals that while the new generation of models are better at solving more instances, they are not necessarily better at solving tasks that require sustained, long-form engagement. ConDABench is an avenue for model builders to measure progress towards truly collaborative models that can complete complex interactive tasks.
title ConDABench: Interactive Evaluation of Language Models for Data Analysis
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.13835