TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yongyao, Miao, Ziqi, Yang, Lu, Jia, Haonan, Yan, Wenting, Qian, Chen, Li, Lijun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915794347098112
author Wang, Yongyao
Miao, Ziqi
Yang, Lu
Jia, Haonan
Yan, Wenting
Qian, Chen
Li, Lijun
author_facet Wang, Yongyao
Miao, Ziqi
Yang, Lu
Jia, Haonan
Yan, Wenting
Qian, Chen
Li, Lijun
contents Tabular prediction can benefit from in-table rows as few-shot evidence, yet existing tabular models typically perform instance-wise inference and LLM-based prompting is often brittle. Models do not consistently leverage relevant rows, and noisy context can degrade performance. To address this challenge, we propose TabSieve, a select-then-predict framework that makes evidence usage explicit and auditable. Given a table and a query row, TabSieve first selects a small set of informative rows as evidence and then predicts the missing target conditioned on the selected evidence. To enable this capability, we construct TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables using a strong teacher model with strict filtering. Furthermore, we introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards, and stabilizes mixed regression and classification training via dynamic task-advantage balancing. Experiments on a held-out benchmark of 75 classification and 52 regression tables show that TabSieve consistently improves performance across shot budgets, with average gains of 2.92% on classification and 4.45% on regression over the second-best baseline. Further analysis indicates that TabSieve concentrates more attention on the selected evidence, which improves robustness to noisy context.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11700
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction
Wang, Yongyao
Miao, Ziqi
Yang, Lu
Jia, Haonan
Yan, Wenting
Qian, Chen
Li, Lijun
Machine Learning
Artificial Intelligence
Tabular prediction can benefit from in-table rows as few-shot evidence, yet existing tabular models typically perform instance-wise inference and LLM-based prompting is often brittle. Models do not consistently leverage relevant rows, and noisy context can degrade performance. To address this challenge, we propose TabSieve, a select-then-predict framework that makes evidence usage explicit and auditable. Given a table and a query row, TabSieve first selects a small set of informative rows as evidence and then predicts the missing target conditioned on the selected evidence. To enable this capability, we construct TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables using a strong teacher model with strict filtering. Furthermore, we introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards, and stabilizes mixed regression and classification training via dynamic task-advantage balancing. Experiments on a held-out benchmark of 75 classification and 52 regression tables show that TabSieve consistently improves performance across shot budgets, with average gains of 2.92% on classification and 4.45% on regression over the second-best baseline. Further analysis indicates that TabSieve concentrates more attention on the selected evidence, which improves robustness to noisy context.
title TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.11700