LLM/Agent-as-Data-Analyst: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zirui, Wang, Weizheng, Zhou, Zihang, Jiao, Yang, Xu, Bangrui, Niu, Boyu, Zhou, Dayou, Zhou, Xuanhe, Li, Guoliang, He, Yeye, Zhou, Wei, Song, Yitong, Tan, Cheng, Yang, Xue, Liu, Chunwei, Wang, Bin, He, Conghui, Wang, Xiaoyang, Wu, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915578606780416
author Tang, Zirui
Wang, Weizheng
Zhou, Zihang
Jiao, Yang
Xu, Bangrui
Niu, Boyu
Zhou, Dayou
Zhou, Xuanhe
Li, Guoliang
He, Yeye
Zhou, Wei
Song, Yitong
Tan, Cheng
Yang, Xue
Liu, Chunwei
Wang, Bin
He, Conghui
Wang, Xiaoyang
Wu, Fan
author_facet Tang, Zirui
Wang, Weizheng
Zhou, Zihang
Jiao, Yang
Xu, Bangrui
Niu, Boyu
Zhou, Dayou
Zhou, Xuanhe
Li, Guoliang
He, Yeye
Zhou, Wei
Song, Yitong
Tan, Cheng
Yang, Xue
Liu, Chunwei
Wang, Bin
He, Conghui
Wang, Xiaoyang
Wu, Fan
contents Large language models (LLMs) and agent techniques have brought a fundamental shift in the functionality and development paradigm of data analysis tasks (a.k.a LLM/Agent-as-Data-Analyst), demonstrating substantial impact across both academia and industry. In comparison with traditional rule or small-model based approaches, (agentic) LLMs enable complex data understanding, natural language interfaces, semantic analysis functions, and autonomous pipeline orchestration. From a modality perspective, we review LLM-based techniques for (i) structured data (e.g., NL2SQL, NL2GQL, ModelQA), (ii) semi-structured data (e.g., markup languages understanding, semi-structured table question answering), (iii) unstructured data (e.g., chart understanding, text/image document understanding), and (iv) heterogeneous data (e.g., data retrieval and modality alignment in data lakes). The technical evolution further distills four key design goals for intelligent data analysis agents, namely semantic-aware design, autonomous pipelines, tool-augmented workflows, and support for open-world tasks. Finally, we outline the remaining challenges and propose several insights and practical directions for advancing LLM/Agent-powered data analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23988
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM/Agent-as-Data-Analyst: A Survey
Tang, Zirui
Wang, Weizheng
Zhou, Zihang
Jiao, Yang
Xu, Bangrui
Niu, Boyu
Zhou, Dayou
Zhou, Xuanhe
Li, Guoliang
He, Yeye
Zhou, Wei
Song, Yitong
Tan, Cheng
Yang, Xue
Liu, Chunwei
Wang, Bin
He, Conghui
Wang, Xiaoyang
Wu, Fan
Artificial Intelligence
Databases
Large language models (LLMs) and agent techniques have brought a fundamental shift in the functionality and development paradigm of data analysis tasks (a.k.a LLM/Agent-as-Data-Analyst), demonstrating substantial impact across both academia and industry. In comparison with traditional rule or small-model based approaches, (agentic) LLMs enable complex data understanding, natural language interfaces, semantic analysis functions, and autonomous pipeline orchestration. From a modality perspective, we review LLM-based techniques for (i) structured data (e.g., NL2SQL, NL2GQL, ModelQA), (ii) semi-structured data (e.g., markup languages understanding, semi-structured table question answering), (iii) unstructured data (e.g., chart understanding, text/image document understanding), and (iv) heterogeneous data (e.g., data retrieval and modality alignment in data lakes). The technical evolution further distills four key design goals for intelligent data analysis agents, namely semantic-aware design, autonomous pipelines, tool-augmented workflows, and support for open-world tasks. Finally, we outline the remaining challenges and propose several insights and practical directions for advancing LLM/Agent-powered data analysis.
title LLM/Agent-as-Data-Analyst: A Survey
topic Artificial Intelligence
Databases
url https://arxiv.org/abs/2509.23988