Saved in:
Bibliographic Details
Main Authors: Łabędzki, Rafał, Miziuła, Patryk, Rutkowski, Hubert, Betlewski, Szymon, Depta, Cezary, Janowski, Szymon, Kochanowicz, Jarosław, Milczek, Jan Kanty
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2606.00051
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914619243626496
author Łabędzki, Rafał
Miziuła, Patryk
Rutkowski, Hubert
Betlewski, Szymon
Depta, Cezary
Janowski, Szymon
Kochanowicz, Jarosław
Milczek, Jan Kanty
author_facet Łabędzki, Rafał
Miziuła, Patryk
Rutkowski, Hubert
Betlewski, Szymon
Depta, Cezary
Janowski, Szymon
Kochanowicz, Jarosław
Milczek, Jan Kanty
contents Large Language Models (LLMs) are increasingly used in analytical workflows, but their suitability as exploratory data analysis (EDA) agents in business settings remains uncertain. In practice, a deployable EDA agent must provide not only useful average performance but also sufficient repeatability to support trust in its outputs. We evaluate this requirement in a controlled, business-relevant benchmark built on an agent-based supply chain simulation. The task is to identify supplier-product combinations responsible for low quality and downstream sales loss by reasoning from indirect operational traces rather than from explicit labels. Fifteen model-variant configurations from eight model families were evaluated under four experimental conditions that varied data representation, prompt clarity, and signal strength, with five trajectories per condition. Outputs were scored against deterministic ground truth using the Jaccard index and assessed through a framework that combines mean score (ms), coefficient of variation (CV), exploratory cross-condition significance tests, and Business utility, a risk-adjusted metric that we propose to summarise quality and repeatability in a single operational measure. The results show that most configurations are not reliable enough for autonomous EDA use, even when their average scores appear acceptable. GPT-5.4 with extra-high reasoning effort achieved the strongest overall profile, with an experiment-averaged ms of 0.8748 and an experiment-averaged Business utility of 0.6952, while the next-best configurations lost substantially more utility after variability discounting. Our findings suggest that evaluation of EDA agents should treat average quality, repeatability, and condition sensitivity as complementary dimensions of operational trustworthiness.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00051
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Business Utility of Large Language Models as Exploratory Data Analysis Agents
Łabędzki, Rafał
Miziuła, Patryk
Rutkowski, Hubert
Betlewski, Szymon
Depta, Cezary
Janowski, Szymon
Kochanowicz, Jarosław
Milczek, Jan Kanty
Computers and Society
Artificial Intelligence
Large Language Models (LLMs) are increasingly used in analytical workflows, but their suitability as exploratory data analysis (EDA) agents in business settings remains uncertain. In practice, a deployable EDA agent must provide not only useful average performance but also sufficient repeatability to support trust in its outputs. We evaluate this requirement in a controlled, business-relevant benchmark built on an agent-based supply chain simulation. The task is to identify supplier-product combinations responsible for low quality and downstream sales loss by reasoning from indirect operational traces rather than from explicit labels. Fifteen model-variant configurations from eight model families were evaluated under four experimental conditions that varied data representation, prompt clarity, and signal strength, with five trajectories per condition. Outputs were scored against deterministic ground truth using the Jaccard index and assessed through a framework that combines mean score (ms), coefficient of variation (CV), exploratory cross-condition significance tests, and Business utility, a risk-adjusted metric that we propose to summarise quality and repeatability in a single operational measure. The results show that most configurations are not reliable enough for autonomous EDA use, even when their average scores appear acceptable. GPT-5.4 with extra-high reasoning effort achieved the strongest overall profile, with an experiment-averaged ms of 0.8748 and an experiment-averaged Business utility of 0.6952, while the next-best configurations lost substantially more utility after variability discounting. Our findings suggest that evaluation of EDA agents should treat average quality, repeatability, and condition sensitivity as complementary dimensions of operational trustworthiness.
title Business Utility of Large Language Models as Exploratory Data Analysis Agents
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2606.00051