Enhancing Table Reasoning with Deterministic Table-State Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Kwok, Tung Sum Thomas, Wang, Xinyu, He, Hengzhi, Lin, Xiaofeng, Lu, Peng, Ma, Liheng, Wang, Chunhe, Mak, Chun Ho, Luo, Yuyu, Wu, Ying Nian, Ding, Lei, Cheng, Guang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Table to Cell: Attention for Better Reasoning with TABALIGN
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
DEREC-SIMPRO: unlock Language Model benefits to advance Synthesis in Data Clean Room
by: Kwok, Tung Sum Thomas, et al.
Published: (2024)
by: Kwok, Tung Sum Thomas, et al.
Published: (2024)
GReaTER: Generate Realistic Tabular data after data Enhancement and Reduction
by: Kwok, Tung Sum Thomas, et al.
Published: (2025)
by: Kwok, Tung Sum Thomas, et al.
Published: (2025)
Co-Evolution of Policy and Internal Reward for Language Agents
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Towards High Supervised Learning Utility Training Data Generation: Data Pruning and Column Reordering
by: Kwok, Tung Sum Thomas, et al.
Published: (2025)
by: Kwok, Tung Sum Thomas, et al.
Published: (2025)
Watermarking Generative Tabular Data
by: He, Hengzhi, et al.
Published: (2024)
by: He, Hengzhi, et al.
Published: (2024)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
TableTale: Reviving the Narrative Interplay Between Data Tables and Text in Scientific Papers
by: Wang, Liangwei, et al.
Published: (2026)
by: Wang, Liangwei, et al.
Published: (2026)
VisTR: Visualizations as Representations for Time-series Table Reasoning
by: Hao, Jianing, et al.
Published: (2024)
by: Hao, Jianing, et al.
Published: (2024)
Watermarking Generative Categorical Data
by: Gu, Bochao, et al.
Published: (2024)
by: Gu, Bochao, et al.
Published: (2024)
Golden Ratio Weighting Prevents Model Collapse
by: He, Hengzhi, et al.
Published: (2025)
by: He, Hengzhi, et al.
Published: (2025)
Let the Target Select for Itself: Data Selection via Target-Aligned Paths
by: Yang, Huitao, et al.
Published: (2026)
by: Yang, Huitao, et al.
Published: (2026)
A Probabilistic Perspective on Model Collapse
by: Xu, Shirong, et al.
Published: (2025)
by: Xu, Shirong, et al.
Published: (2025)
Seek and Solve Reasoning for Table Question Answering
by: Jiang, Ruya, et al.
Published: (2024)
by: Jiang, Ruya, et al.
Published: (2024)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
by: Feng, Xiao, et al.
Published: (2026)
by: Feng, Xiao, et al.
Published: (2026)
Building disciplinary literacies in content and language integrated learning By JuliaHüttner, ChristianeDalton‐Puffer (Eds.), New York: Routledge. 2024. xii + pp. 225. (hbk) ISBN: 9781032517292
by: Hengzhi Hu
Published: (2025)
by: Hengzhi Hu
Published: (2025)
First Pediatric Application of Bachmann's Bundle Pacing and Left Bundle Branch Area Pacing for Bi‐Physiologic Conduction System Pacing
by: Hei‐To Leung, et al.
Published: (2026)
by: Hei‐To Leung, et al.
Published: (2026)
Knowledge Reasoning Involving Four Types of Syllogisms
by: Wei, Long, et al.
Published: (2025)
by: Wei, Long, et al.
Published: (2025)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
FBChain: A Blockchain-based Federated Learning Model with Efficiency and Secure Communication
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
Editorial: Revolutionising Steatotic Liver Disease Diagnosis With Phosphatidylethanol
by: Karen Cheuk‐Ying Ho, et al.
Published: (2025)
by: Karen Cheuk‐Ying Ho, et al.
Published: (2025)
Authenticated Contradictions from Desynchronized Provenance and Watermarking
by: Nemecek, Alexander, et al.
Published: (2026)
by: Nemecek, Alexander, et al.
Published: (2026)
Explore the Reasoning Capability of LLMs in the Chess Testbed
by: Wang, Shu, et al.
Published: (2024)
by: Wang, Shu, et al.
Published: (2024)
$\text{R}^2\text{R}$: A Route-to-Rerank Post-Training Framework for Multi-Domain Decoder-Only Rerankers
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
Online Deterministic Minimum Cost Bipartite Matching with Delays on a Line
by: Kuo, Tung-Wei
Published: (2024)
by: Kuo, Tung-Wei
Published: (2024)
Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
by: Lei, Fangyu, et al.
Published: (2025)
by: Lei, Fangyu, et al.
Published: (2025)
Adaptive Dual-domain Learning for Underwater Image Enhancement
by: Peng, Lingtao, et al.
Published: (2025)
by: Peng, Lingtao, et al.
Published: (2025)
From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Training Retrieval-Augmented Generation Agents
by: Li, Muzhi, et al.
Published: (2025)
by: Li, Muzhi, et al.
Published: (2025)
Stateful Reasoning via Insight Replay
by: Lei, Bin, et al.
Published: (2026)
by: Lei, Bin, et al.
Published: (2026)
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
by: Shi, Jingze, et al.
Published: (2026)
by: Shi, Jingze, et al.
Published: (2026)
When Stochastic Rewards Reduce to Deterministic Rewards in Online Bipartite Matching
by: Udwani, Rajan
Published: (2023)
by: Udwani, Rajan
Published: (2023)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025)
by: Luo, Ruilin, et al.
Published: (2025)
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
by: Wang, Xiaokun, et al.
Published: (2025)
by: Wang, Xiaokun, et al.
Published: (2025)
Similar Items
-
From Table to Cell: Attention for Better Reasoning with TABALIGN
by: Kwok, Tung Sum Thomas, et al.
Published: (2026) -
TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering
by: Kwok, Tung Sum Thomas, et al.
Published: (2026) -
DEREC-SIMPRO: unlock Language Model benefits to advance Synthesis in Data Clean Room
by: Kwok, Tung Sum Thomas, et al.
Published: (2024) -
GReaTER: Generate Realistic Tabular data after data Enhancement and Reduction
by: Kwok, Tung Sum Thomas, et al.
Published: (2025) -
Co-Evolution of Policy and Internal Reward for Language Agents
by: Wang, Xinyu, et al.
Published: (2026)