Evaluating Latent Knowledge of Public Tabular Datasets in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Silvestri, Matteo, Veglianti, Fabiano, Giorgi, Flavio, Silvestri, Fabrizio, Tolomei, Gabriele
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915897793314816
author Silvestri, Matteo
Veglianti, Fabiano
Giorgi, Flavio
Silvestri, Fabrizio
Tolomei, Gabriele
author_facet Silvestri, Matteo
Veglianti, Fabiano
Giorgi, Flavio
Silvestri, Fabrizio
Tolomei, Gabriele
contents Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datasets rather than generalization. However, in the context of tabular data, this problem is largely unexplored. Existing approaches primarily rely on memorization tests, which are too coarse to detect contamination. In contrast, we propose a framework for assessing contamination in tabular datasets by generating controlled queries and performing comparative evaluation. Given a dataset, we craft multiple-choice aligned queries that preserve task structure while allowing systematic transformations of the underlying data. These transformations are designed to selectively disrupt dataset information while preserving partial knowledge, enabling us to isolate performance attributable to contamination. We complement this setup with non-neural baselines that provide reference performance, and we introduce a statistical testing procedure to formally detect significant deviations indicative of contamination. Empirical results on eight widely used tabular datasets reveal clear evidence of contamination in four cases. These findings suggest that performance on downstream tasks involving such datasets may be substantially inflated, raising concerns about the reliability of current evaluation practices.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20351
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Latent Knowledge of Public Tabular Datasets in Large Language Models
Silvestri, Matteo
Veglianti, Fabiano
Giorgi, Flavio
Silvestri, Fabrizio
Tolomei, Gabriele
Computation and Language
Artificial Intelligence
Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datasets rather than generalization. However, in the context of tabular data, this problem is largely unexplored. Existing approaches primarily rely on memorization tests, which are too coarse to detect contamination. In contrast, we propose a framework for assessing contamination in tabular datasets by generating controlled queries and performing comparative evaluation. Given a dataset, we craft multiple-choice aligned queries that preserve task structure while allowing systematic transformations of the underlying data. These transformations are designed to selectively disrupt dataset information while preserving partial knowledge, enabling us to isolate performance attributable to contamination. We complement this setup with non-neural baselines that provide reference performance, and we introduce a statistical testing procedure to formally detect significant deviations indicative of contamination. Empirical results on eight widely used tabular datasets reveal clear evidence of contamination in four cases. These findings suggest that performance on downstream tasks involving such datasets may be substantially inflated, raising concerns about the reliability of current evaluation practices.
title Evaluating Latent Knowledge of Public Tabular Datasets in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.20351