FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Onderková, Kristýna, Plátek, Ondřej, Kasner, Zdeněk, Dušek, Ondřej
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911211941003264
author Onderková, Kristýna
Plátek, Ondřej
Kasner, Zdeněk
Dušek, Ondřej
author_facet Onderková, Kristýna
Plátek, Ondřej
Kasner, Zdeněk
Dušek, Ondřej
contents Table-to-text generation (insight generation from tables) is a challenging task that requires precision in analyzing the data. In addition, the evaluation of existing benchmarks is affected by contamination of Large Language Model (LLM) training data as well as domain imbalance. We introduce FreshTab, an on-the-fly table-to-text benchmark generation from Wikipedia, to combat the LLM data contamination problem and enable domain-sensitive evaluation. While non-English table-to-text datasets are limited, FreshTab collects datasets in different languages on demand (we experiment with German, Russian and French in addition to English). We find that insights generated by LLMs from recent tables collected by our method appear clearly worse by automatic metrics, but this does not translate into LLM and human evaluations. Domain effects are visible in all evaluations, showing that a~domain-balanced benchmark is more challenging.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13598
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
Onderková, Kristýna
Plátek, Ondřej
Kasner, Zdeněk
Dušek, Ondřej
Computation and Language
Table-to-text generation (insight generation from tables) is a challenging task that requires precision in analyzing the data. In addition, the evaluation of existing benchmarks is affected by contamination of Large Language Model (LLM) training data as well as domain imbalance. We introduce FreshTab, an on-the-fly table-to-text benchmark generation from Wikipedia, to combat the LLM data contamination problem and enable domain-sensitive evaluation. While non-English table-to-text datasets are limited, FreshTab collects datasets in different languages on demand (we experiment with German, Russian and French in addition to English). We find that insights generated by LLMs from recent tables collected by our method appear clearly worse by automatic metrics, but this does not translate into LLM and human evaluations. Domain effects are visible in all evaluations, showing that a~domain-balanced benchmark is more challenging.
title FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
topic Computation and Language
url https://arxiv.org/abs/2510.13598