p2-TQA: A Process-based Preference Learning Framework for Self-Improving Table Question Answering Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Wei, Mesgar, Mohsen, Adel, Heike, Friedrich, Annemarie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908641405173760
author Zhou, Wei
Mesgar, Mohsen
Adel, Heike
Friedrich, Annemarie
author_facet Zhou, Wei
Mesgar, Mohsen
Adel, Heike
Friedrich, Annemarie
contents Table question answering (TQA) focuses on answering questions based on tabular data. Developing TQA systems targets effective interaction with tabular data for tasks such as cell retrieval and data analysis. While recent work has leveraged fine-tuning to improve TQA systems, existing approaches often under-utilize available data and neglect the potential of post-training for further gains. In this work, we introduce p2-TQA, a process-based preference learning framework for TQA post-training. p2-TQA automatically constructs process-based preference data via a table-specific pipeline, eliminating the need for manual or costly data collection. It then optimizes models through contrastive learning on the collected data. Experiments show that p2-TQA effectively improves TQA models by up to 5% on in-domain datasets and 2.4% on out-of-domain datasets with only 8,000 training instances. Furthermore, models enhanced with p2-TQA achieve competitive results against larger, more complex state-of-the-art TQA systems, while maintaining up to five times higher efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle p2-TQA: A Process-based Preference Learning Framework for Self-Improving Table Question Answering Models
Zhou, Wei
Mesgar, Mohsen
Adel, Heike
Friedrich, Annemarie
Computation and Language
Table question answering (TQA) focuses on answering questions based on tabular data. Developing TQA systems targets effective interaction with tabular data for tasks such as cell retrieval and data analysis. While recent work has leveraged fine-tuning to improve TQA systems, existing approaches often under-utilize available data and neglect the potential of post-training for further gains. In this work, we introduce p2-TQA, a process-based preference learning framework for TQA post-training. p2-TQA automatically constructs process-based preference data via a table-specific pipeline, eliminating the need for manual or costly data collection. It then optimizes models through contrastive learning on the collected data. Experiments show that p2-TQA effectively improves TQA models by up to 5% on in-domain datasets and 2.4% on out-of-domain datasets with only 8,000 training instances. Furthermore, models enhanced with p2-TQA achieve competitive results against larger, more complex state-of-the-art TQA systems, while maintaining up to five times higher efficiency.
title p2-TQA: A Process-based Preference Learning Framework for Self-Improving Table Question Answering Models
topic Computation and Language
url https://arxiv.org/abs/2505.17565