Probing How Scalable Table Data Enhances General Long-Context Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915883088084992 |
|---|---|
| author | Xie, Huaibing Zhao, Guoliang Liu, Yang Dou, Shihan Huang, Siming Xiao, Yanling Wang, Shaolei Liu, Yiting Zhang, Cheng Liu, Shaofan Zhou, Pluto |
| author_facet | Xie, Huaibing Zhao, Guoliang Liu, Yang Dou, Shihan Huang, Siming Xiao, Yanling Wang, Shaolei Liu, Yiting Zhang, Cheng Liu, Shaofan Zhou, Pluto |
| contents | As real-world tasks grow increasingly complex, long-context reasoning has become a core capability for Large Language Models (LLMs). However, few studies explore which data types are effective for long-context reasoning and why. We find that structured table data with periodic structures shows strong potential for long-context reasoning. Motivated by this observation, we mathematically analyze tabular dependency structures using mutual information, revealing periodic non-vanishing dependencies in table data. Furthermore, we systematically analyze the capabilities of structured table data, conduct relevant scaling experiments, and validate its underlying mechanisms for enhancing long-context reasoning, yielding several meaningful insights. Leveraging these insights, we propose a simple yet scalable pipeline(TableLong) for synthesizing high-quality, diverse, and verifiable structured table data to boost long-context reasoning via RL. Extensive experimental results demonstrate that table data significantly enhances the long-context reasoning capability of LLMs across multiple long-context benchmarks (+8.24\% on average), and even improves performance on out-of-domain benchmarks (+8.06\% on average). We hope that our insights provide practical guidance for effective post-training data to enhance long-context reasoning in LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_21719 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Probing How Scalable Table Data Enhances General Long-Context Reasoning Xie, Huaibing Zhao, Guoliang Liu, Yang Dou, Shihan Huang, Siming Xiao, Yanling Wang, Shaolei Liu, Yiting Zhang, Cheng Liu, Shaofan Zhou, Pluto Computation and Language As real-world tasks grow increasingly complex, long-context reasoning has become a core capability for Large Language Models (LLMs). However, few studies explore which data types are effective for long-context reasoning and why. We find that structured table data with periodic structures shows strong potential for long-context reasoning. Motivated by this observation, we mathematically analyze tabular dependency structures using mutual information, revealing periodic non-vanishing dependencies in table data. Furthermore, we systematically analyze the capabilities of structured table data, conduct relevant scaling experiments, and validate its underlying mechanisms for enhancing long-context reasoning, yielding several meaningful insights. Leveraging these insights, we propose a simple yet scalable pipeline(TableLong) for synthesizing high-quality, diverse, and verifiable structured table data to boost long-context reasoning via RL. Extensive experimental results demonstrate that table data significantly enhances the long-context reasoning capability of LLMs across multiple long-context benchmarks (+8.24\% on average), and even improves performance on out-of-domain benchmarks (+8.06\% on average). We hope that our insights provide practical guidance for effective post-training data to enhance long-context reasoning in LLMs. |
| title | Probing How Scalable Table Data Enhances General Long-Context Reasoning |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2603.21719 |