Self-Correction Distillation for Structured Data Question Answering

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Yushan, Zhang, Wen, Jin, Long, Sun, Mengshu, Zhong, Ling, Liu, Zhiqiang, Li, Juan, Liang, Lei, Long, Chong, Deng, Chao, Feng, Junlan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918205263446016
author Zhu, Yushan
Zhang, Wen
Jin, Long
Sun, Mengshu
Zhong, Ling
Liu, Zhiqiang
Li, Juan
Liang, Lei
Long, Chong
Deng, Chao
Feng, Junlan
author_facet Zhu, Yushan
Zhang, Wen
Jin, Long
Sun, Mengshu
Zhong, Ling
Liu, Zhiqiang
Li, Juan
Liang, Lei
Long, Chong
Deng, Chao
Feng, Junlan
contents Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07998
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Correction Distillation for Structured Data Question Answering
Zhu, Yushan
Zhang, Wen
Jin, Long
Sun, Mengshu
Zhong, Ling
Liu, Zhiqiang
Li, Juan
Liang, Lei
Long, Chong
Deng, Chao
Feng, Junlan
Computation and Language
Artificial Intelligence
Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets.
title Self-Correction Distillation for Structured Data Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.07998