JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909947922481152 |
|---|---|
| author | Chi, Ce Wang, Xing Wang, Zhendong Liu, Xiaofan Li, Ce Song, Zhiyan Zhao, Chen Yang, Kexin Shi, Boshen Yang, Jingjing Deng, Chao Feng, Junlan |
| author_facet | Chi, Ce Wang, Xing Wang, Zhendong Liu, Xiaofan Li, Ce Song, Zhiyan Zhao, Chen Yang, Kexin Shi, Boshen Yang, Jingjing Deng, Chao Feng, Junlan |
| contents | In this work, we present JT-DA-8B (JiuTian Data Analyst 8B), a specialized large language model designed for complex table reasoning tasks across diverse real-world scenarios. To address the lack of high-quality supervision in tabular reasoning scenarios, we construct a comprehensive and diverse training corpus with 34 well-defined table reasoning tasks, by aggregating 29 public table QA datasets and 3 million tables. An automatic pipeline is proposed to generate realistic multi-step analytical tasks involving reasoning patterns. The model is trained upon open-source JT-Coder-8B model, an 8B-parameter decoder-only foundation model trained from scratch. In the training stage, we leverage LLM-based scoring and workflow-aligned filtering to distill high-quality, table-centric data. Both supervised fine-tuning (SFT) and Reinforcement learning (RL) are adopted to optimize our model. Afterwards, a four-stage table reasoning workflow is proposed, including table preprocessing, table sensing, tool-integrated reasoning, and prompt engineering, to improve model interpretability and execution accuracy. Experimental results show that JT-DA-8B achieves strong performance in various table reasoning tasks, demonstrating the effectiveness of data-centric generation and workflow-driven optimization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_06859 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models Chi, Ce Wang, Xing Wang, Zhendong Liu, Xiaofan Li, Ce Song, Zhiyan Zhao, Chen Yang, Kexin Shi, Boshen Yang, Jingjing Deng, Chao Feng, Junlan Artificial Intelligence In this work, we present JT-DA-8B (JiuTian Data Analyst 8B), a specialized large language model designed for complex table reasoning tasks across diverse real-world scenarios. To address the lack of high-quality supervision in tabular reasoning scenarios, we construct a comprehensive and diverse training corpus with 34 well-defined table reasoning tasks, by aggregating 29 public table QA datasets and 3 million tables. An automatic pipeline is proposed to generate realistic multi-step analytical tasks involving reasoning patterns. The model is trained upon open-source JT-Coder-8B model, an 8B-parameter decoder-only foundation model trained from scratch. In the training stage, we leverage LLM-based scoring and workflow-aligned filtering to distill high-quality, table-centric data. Both supervised fine-tuning (SFT) and Reinforcement learning (RL) are adopted to optimize our model. Afterwards, a four-stage table reasoning workflow is proposed, including table preprocessing, table sensing, tool-integrated reasoning, and prompt engineering, to improve model interpretability and execution accuracy. Experimental results show that JT-DA-8B achieves strong performance in various table reasoning tasks, demonstrating the effectiveness of data-centric generation and workflow-driven optimization. |
| title | JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2512.06859 |