Reinforcing Numerical Reasoning in LLMs for Tabular Prediction via Structural Priors

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Pengxiang, Gao, Zihao, Lian, Wanchen, Chen, Jintai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910046862966784
author Cai, Pengxiang
Gao, Zihao
Lian, Wanchen
Chen, Jintai
author_facet Cai, Pengxiang
Gao, Zihao
Lian, Wanchen
Chen, Jintai
contents Tabular prediction traditionally relies on gradient-boosted decision trees and deep learning models, which excel in specific tasks but lack interpretability and transferability. Reasoning large language models (LLMs) promise cross-task adaptability with transparent reasoning traces, yet their potential for tabular data remains unrealized. To bridge this gap, we propose a reasoning framework centered on Permutation Relative Policy Optimization (PRPO), a reinforcement learning method that encodes column-permutation invariance as a structural prior. By estimating advantages across label-preserving permutations, PRPO transforms sparse rewards into dense signals, activating latent numerical reasoning capabilities of LLMs with limited supervision. Extensive experiments show that our method matches fully supervised baselines and dominates in zero-shot settings, performing on par with 32-shot strong baselines. Remarkably, our 8B model significantly outperforms much larger LLMs, achieving up to a 53.17% improvement over DeepSeek-R1 (685B).
format Preprint
id arxiv_https___arxiv_org_abs_2510_17385
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcing Numerical Reasoning in LLMs for Tabular Prediction via Structural Priors
Cai, Pengxiang
Gao, Zihao
Lian, Wanchen
Chen, Jintai
Machine Learning
Artificial Intelligence
Tabular prediction traditionally relies on gradient-boosted decision trees and deep learning models, which excel in specific tasks but lack interpretability and transferability. Reasoning large language models (LLMs) promise cross-task adaptability with transparent reasoning traces, yet their potential for tabular data remains unrealized. To bridge this gap, we propose a reasoning framework centered on Permutation Relative Policy Optimization (PRPO), a reinforcement learning method that encodes column-permutation invariance as a structural prior. By estimating advantages across label-preserving permutations, PRPO transforms sparse rewards into dense signals, activating latent numerical reasoning capabilities of LLMs with limited supervision. Extensive experiments show that our method matches fully supervised baselines and dominates in zero-shot settings, performing on par with 32-shot strong baselines. Remarkably, our 8B model significantly outperforms much larger LLMs, achieving up to a 53.17% improvement over DeepSeek-R1 (685B).
title Reinforcing Numerical Reasoning in LLMs for Tabular Prediction via Structural Priors
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.17385