Effective Distillation of Table-based Reasoning Ability from LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Bohao, Tang, Chen, Zhao, Kun, Xiao, Chenghao, Lin, Chenghua
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917621312520192
author Yang, Bohao
Tang, Chen
Zhao, Kun
Xiao, Chenghao
Lin, Chenghua
author_facet Yang, Bohao
Tang, Chen
Zhao, Kun
Xiao, Chenghao
Lin, Chenghua
contents Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their enormous parameter size and extremely high requirements for compute power pose challenges for their practical deployment. Recent research has revealed that specific capabilities of LLMs, such as numerical reasoning, can be transferred to smaller models through distillation. Some studies explore the potential of leveraging LLMs to perform table-based reasoning. However, there has been no prior work focusing on table reasoning skills in smaller models specifically tailored for scientific table-to-text generation tasks. In this paper, we propose a novel table-based reasoning distillation approach, with the aim of distilling LLMs into tailored smaller models. Our experimental results have shown that a 220 million parameter model (Flan-T5-base) fine-tuned using distilled data, not only achieves a significant improvement compared to traditionally fine-tuned baselines, but also surpasses specific LLMs on a scientific table-to-text generation dataset. Our code is available at https://github.com/Bernard-Yang/DistillTableCoT.
format Preprint
id arxiv_https___arxiv_org_abs_2309_13182
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Effective Distillation of Table-based Reasoning Ability from LLMs
Yang, Bohao
Tang, Chen
Zhao, Kun
Xiao, Chenghao
Lin, Chenghua
Computation and Language
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their enormous parameter size and extremely high requirements for compute power pose challenges for their practical deployment. Recent research has revealed that specific capabilities of LLMs, such as numerical reasoning, can be transferred to smaller models through distillation. Some studies explore the potential of leveraging LLMs to perform table-based reasoning. However, there has been no prior work focusing on table reasoning skills in smaller models specifically tailored for scientific table-to-text generation tasks. In this paper, we propose a novel table-based reasoning distillation approach, with the aim of distilling LLMs into tailored smaller models. Our experimental results have shown that a 220 million parameter model (Flan-T5-base) fine-tuned using distilled data, not only achieves a significant improvement compared to traditionally fine-tuned baselines, but also surpasses specific LLMs on a scientific table-to-text generation dataset. Our code is available at https://github.com/Bernard-Yang/DistillTableCoT.
title Effective Distillation of Table-based Reasoning Ability from LLMs
topic Computation and Language
url https://arxiv.org/abs/2309.13182