GLEAN: Grounded Lightweight Evaluation Anchors for Contamination-Aware Tabular Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Wang, Qizhi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910038222700544
author Wang, Qizhi
author_facet Wang, Qizhi
contents Tabular reasoning benchmarks mix semantic inference, numerical computation, and brittle table formatting, yet evaluations for small models remain vulnerable to contamination, dataset artifacts, and retrieval failures. We propose GLEAN, a lightweight evaluation protocol that integrates contamination-aware probes, weak-supervision governance, retrieval-reasoning diagnostics, and structured error attribution under tight hardware constraints. We evaluate across TabFact, WTQ via Squall, TableBench, RobuT, and SciTab under a 16GB GPU budget. Using Squall gold SQL as an executable anchor (95.2% execution), GLEAN assigns a deterministic error taxonomy (L0-L4 plus L0.5 context miss) and reveals a stable error-mode separation: TAPEX errors skew toward grounding (L3) while TAPAS errors skew toward hallucination/abstention (L2/L0). We validate evidence-row heuristics against SQL-derived rows on simple queries (0.62 precision / 0.71 recall; hybrid recall 0.81) and show that retrieval Recall@K can saturate even when end-to-end EM/F1 remains limited, motivating attribution beyond raw recall. We release a modular framework with audits and sensitivity checks to make small-model tabular evaluation more contamination-aware and diagnostic.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02212
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GLEAN: Grounded Lightweight Evaluation Anchors for Contamination-Aware Tabular Reasoning
Wang, Qizhi
Databases
Artificial Intelligence
68T01 (Primary) 68T05, 68T50, 68P20 (Secondary)
I.2.7; H.3.3; H.2.8; I.2.1
Tabular reasoning benchmarks mix semantic inference, numerical computation, and brittle table formatting, yet evaluations for small models remain vulnerable to contamination, dataset artifacts, and retrieval failures. We propose GLEAN, a lightweight evaluation protocol that integrates contamination-aware probes, weak-supervision governance, retrieval-reasoning diagnostics, and structured error attribution under tight hardware constraints. We evaluate across TabFact, WTQ via Squall, TableBench, RobuT, and SciTab under a 16GB GPU budget. Using Squall gold SQL as an executable anchor (95.2% execution), GLEAN assigns a deterministic error taxonomy (L0-L4 plus L0.5 context miss) and reveals a stable error-mode separation: TAPEX errors skew toward grounding (L3) while TAPAS errors skew toward hallucination/abstention (L2/L0). We validate evidence-row heuristics against SQL-derived rows on simple queries (0.62 precision / 0.71 recall; hybrid recall 0.81) and show that retrieval Recall@K can saturate even when end-to-end EM/F1 remains limited, motivating attribution beyond raw recall. We release a modular framework with audits and sensitivity checks to make small-model tabular evaluation more contamination-aware and diagnostic.
title GLEAN: Grounded Lightweight Evaluation Anchors for Contamination-Aware Tabular Reasoning
topic Databases
Artificial Intelligence
68T01 (Primary) 68T05, 68T50, 68P20 (Secondary)
I.2.7; H.3.3; H.2.8; I.2.1
url https://arxiv.org/abs/2603.02212