FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Ziting, Zhang, Shize, Yuan, Haitao, Zhu, Jinwei, Dong, Wei, Cong, Gao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911730125242368
author Wang, Ziting
Zhang, Shize
Yuan, Haitao
Zhu, Jinwei
Dong, Wei
Cong, Gao
author_facet Wang, Ziting
Zhang, Shize
Yuan, Haitao
Zhu, Jinwei
Dong, Wei
Cong, Gao
contents The growing demand for data-driven decision-making has created an urgent need for data agents that can reason over heterogeneous data (databases, documents, web content, images, videos, and audio) to answer complex analytical queries. However, evaluating such agents remains challenging: existing benchmarks often focus on isolated agent capabilities or limited data modalities, lacking comprehensive coverage of heterogeneous data and rigorous evaluation across diverse data agent architectures. To address these challenges, we present FDABench, a benchmark for evaluating data agents' reasoning ability over heterogeneous data in analytical scenarios. Our contributions are threefold: (1) A comprehensive benchmark of 2,007 tasks spanning six data modalities with a unified, multi-granularity evaluation framework. (2) We design PUDDING, an agentic dataset construction framework that leverages LLM generation with iterative expert validation for reliable and scalable benchmark construction. (3) Extensive experiments across diverse data agent architectures, including general analytical agents, semantic operator frameworks, and RAG-based methods, revealing key insights and guidelines for future data agent development. Our data and source code are released at https://github.com/fdabench/FDAbench.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
Wang, Ziting
Zhang, Shize
Yuan, Haitao
Zhu, Jinwei
Dong, Wei
Cong, Gao
Databases
The growing demand for data-driven decision-making has created an urgent need for data agents that can reason over heterogeneous data (databases, documents, web content, images, videos, and audio) to answer complex analytical queries. However, evaluating such agents remains challenging: existing benchmarks often focus on isolated agent capabilities or limited data modalities, lacking comprehensive coverage of heterogeneous data and rigorous evaluation across diverse data agent architectures. To address these challenges, we present FDABench, a benchmark for evaluating data agents' reasoning ability over heterogeneous data in analytical scenarios. Our contributions are threefold: (1) A comprehensive benchmark of 2,007 tasks spanning six data modalities with a unified, multi-granularity evaluation framework. (2) We design PUDDING, an agentic dataset construction framework that leverages LLM generation with iterative expert validation for reliable and scalable benchmark construction. (3) Extensive experiments across diverse data agent architectures, including general analytical agents, semantic operator frameworks, and RAG-based methods, revealing key insights and guidelines for future data agent development. Our data and source code are released at https://github.com/fdabench/FDAbench.
title FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
topic Databases
url https://arxiv.org/abs/2509.02473