Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shiqiang, Woisetschläger, Herbert, Jacobsen, Hans Arno, Ji, Mingyue
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913143496638464
author Wang, Shiqiang
Woisetschläger, Herbert
Jacobsen, Hans Arno
Ji, Mingyue
author_facet Wang, Shiqiang
Woisetschläger, Herbert
Jacobsen, Hans Arno
Ji, Mingyue
contents Data is fundamental to large language models (LLMs). However, understanding of what makes certain data useful for different stages of an LLM workflow, including training, tuning, alignment, in-context learning, etc., and why, remains an open question. Current approaches rely heavily on extensive experimentation with large public datasets to obtain empirical heuristics for data filtering and dataset construction. These approaches are compute intensive and lack a principled way of understanding the essence of how specific data characteristics drive LLM behavior. In this position paper, we advocate for the need of developing systematic methodologies for generating synthetic sequences from appropriately defined random processes, with the goal that these sequences can reveal useful characteristics when they are used in one or multiple stages of the LLM workflow. We refer to such sequences as data probes. By observing LLM behavior on data probes, researchers can systematically conduct studies on how data characteristics influence model performance, generalization, and robustness. The probing sequences exhibit statistical properties that can be viewed using theoretical concepts, such as typical sets, which are generalized to describe the behaviors of LLMs. This data-probe approach provides a pathway for uncovering foundational insights into the role of data in LLM training and inference, beyond empirical heuristics.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18801
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
Wang, Shiqiang
Woisetschläger, Herbert
Jacobsen, Hans Arno
Ji, Mingyue
Artificial Intelligence
Information Retrieval
Machine Learning
Data is fundamental to large language models (LLMs). However, understanding of what makes certain data useful for different stages of an LLM workflow, including training, tuning, alignment, in-context learning, etc., and why, remains an open question. Current approaches rely heavily on extensive experimentation with large public datasets to obtain empirical heuristics for data filtering and dataset construction. These approaches are compute intensive and lack a principled way of understanding the essence of how specific data characteristics drive LLM behavior. In this position paper, we advocate for the need of developing systematic methodologies for generating synthetic sequences from appropriately defined random processes, with the goal that these sequences can reveal useful characteristics when they are used in one or multiple stages of the LLM workflow. We refer to such sequences as data probes. By observing LLM behavior on data probes, researchers can systematically conduct studies on how data characteristics influence model performance, generalization, and robustness. The probing sequences exhibit statistical properties that can be viewed using theoretical concepts, such as typical sets, which are generalized to describe the behaviors of LLMs. This data-probe approach provides a pathway for uncovering foundational insights into the role of data in LLM training and inference, beyond empirical heuristics.
title Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
topic Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2605.18801