Self-Correction Distillation for Structured Data Question Answering
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918205263446016 |
|---|---|
| author | Zhu, Yushan Zhang, Wen Jin, Long Sun, Mengshu Zhong, Ling Liu, Zhiqiang Li, Juan Liang, Lei Long, Chong Deng, Chao Feng, Junlan |
| author_facet | Zhu, Yushan Zhang, Wen Jin, Long Sun, Mengshu Zhong, Ling Liu, Zhiqiang Li, Juan Liang, Lei Long, Chong Deng, Chao Feng, Junlan |
| contents | Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_07998 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Self-Correction Distillation for Structured Data Question Answering Zhu, Yushan Zhang, Wen Jin, Long Sun, Mengshu Zhong, Ling Liu, Zhiqiang Li, Juan Liang, Lei Long, Chong Deng, Chao Feng, Junlan Computation and Language Artificial Intelligence Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets. |
| title | Self-Correction Distillation for Structured Data Question Answering |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2511.07998 |