Learning to Maximize Mutual Information for Chain-of-Thought Distillation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Xin, Huang, Hanxian, Gao, Yanjun, Wang, Yi, Zhao, Jishen, Ding, Ke
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913383518830592
author Chen, Xin
Huang, Hanxian
Gao, Yanjun
Wang, Yi
Zhao, Jishen
Ding, Ke
author_facet Chen, Xin
Huang, Hanxian
Gao, Yanjun
Wang, Yi
Zhao, Jishen
Ding, Ke
contents Knowledge distillation, the technique of transferring knowledge from large, complex models to smaller ones, marks a pivotal step towards efficient AI deployment. Distilling Step-by-Step~(DSS), a novel method utilizing chain-of-thought~(CoT) distillation, has demonstrated promise by imbuing smaller models with the superior reasoning capabilities of their larger counterparts. In DSS, the distilled model acquires the ability to generate rationales and predict labels concurrently through a multi-task learning framework. However, DSS overlooks the intrinsic relationship between the two training tasks, leading to ineffective integration of CoT knowledge with the task of label prediction. To this end, we investigate the mutual relationship of the two tasks from Information Bottleneck perspective and formulate it as maximizing the mutual information of the representation features of the two tasks. We propose a variational approach to solve this optimization problem using a learning-based method. Our experimental results across four datasets demonstrate that our method outperforms the state-of-the-art DSS. Our findings offer insightful guidance for future research on language model distillation as well as applications involving CoT. Codes are available at \url{https://github.com/xinchen9/cot_distillation_ACL2024}.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03348
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Maximize Mutual Information for Chain-of-Thought Distillation
Chen, Xin
Huang, Hanxian
Gao, Yanjun
Wang, Yi
Zhao, Jishen
Ding, Ke
Computation and Language
Artificial Intelligence
Knowledge distillation, the technique of transferring knowledge from large, complex models to smaller ones, marks a pivotal step towards efficient AI deployment. Distilling Step-by-Step~(DSS), a novel method utilizing chain-of-thought~(CoT) distillation, has demonstrated promise by imbuing smaller models with the superior reasoning capabilities of their larger counterparts. In DSS, the distilled model acquires the ability to generate rationales and predict labels concurrently through a multi-task learning framework. However, DSS overlooks the intrinsic relationship between the two training tasks, leading to ineffective integration of CoT knowledge with the task of label prediction. To this end, we investigate the mutual relationship of the two tasks from Information Bottleneck perspective and formulate it as maximizing the mutual information of the representation features of the two tasks. We propose a variational approach to solve this optimization problem using a learning-based method. Our experimental results across four datasets demonstrate that our method outperforms the state-of-the-art DSS. Our findings offer insightful guidance for future research on language model distillation as well as applications involving CoT. Codes are available at \url{https://github.com/xinchen9/cot_distillation_ACL2024}.
title Learning to Maximize Mutual Information for Chain-of-Thought Distillation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.03348