Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Jooyoung, Yang, Fan, Tran, Thanh, Hu, Qian, Barut, Emre, Chang, Kai-Wei, Su, Chengwei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914741144780800
author Lee, Jooyoung
Yang, Fan
Tran, Thanh
Hu, Qian
Barut, Emre
Chang, Kai-Wei
Su, Chengwei
author_facet Lee, Jooyoung
Yang, Fan
Tran, Thanh
Hu, Qian
Barut, Emre
Chang, Kai-Wei
Su, Chengwei
contents We introduce a novel framework, LM-Guided CoT, that leverages a lightweight (i.e., <1B) language model (LM) for guiding a black-box large (i.e., >10B) LM in reasoning tasks. Specifically, the lightweight LM first generates a rationale for each input instance. The Frozen large LM is then prompted to predict a task output based on the rationale generated by the lightweight LM. Our approach is resource-efficient in the sense that it only requires training the lightweight LM. We optimize the model through 1) knowledge distillation and 2) reinforcement learning from rationale-oriented and task-oriented reward signals. We assess our method with multi-hop extractive question answering (QA) benchmarks, HotpotQA, and 2WikiMultiHopQA. Experimental results show that our approach outperforms all baselines regarding answer prediction accuracy. We also find that reinforcement learning helps the model to produce higher-quality rationales with improved QA performance.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03414
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought
Lee, Jooyoung
Yang, Fan
Tran, Thanh
Hu, Qian
Barut, Emre
Chang, Kai-Wei
Su, Chengwei
Computation and Language
Artificial Intelligence
We introduce a novel framework, LM-Guided CoT, that leverages a lightweight (i.e., <1B) language model (LM) for guiding a black-box large (i.e., >10B) LM in reasoning tasks. Specifically, the lightweight LM first generates a rationale for each input instance. The Frozen large LM is then prompted to predict a task output based on the rationale generated by the lightweight LM. Our approach is resource-efficient in the sense that it only requires training the lightweight LM. We optimize the model through 1) knowledge distillation and 2) reinforcement learning from rationale-oriented and task-oriented reward signals. We assess our method with multi-hop extractive question answering (QA) benchmarks, HotpotQA, and 2WikiMultiHopQA. Experimental results show that our approach outperforms all baselines regarding answer prediction accuracy. We also find that reinforcement learning helps the model to produce higher-quality rationales with improved QA performance.
title Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.03414