Teaching LLMs to Refine with Tools

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Dian, Zhang, Yuheng, Xu, Jiahao, Liang, Tian, Song, Linfeng, Tu, Zhaopeng, Mi, Haitao, Yu, Dong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912165960613888
author Yu, Dian
Zhang, Yuheng
Xu, Jiahao
Liang, Tian
Song, Linfeng
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
author_facet Yu, Dian
Zhang, Yuheng
Xu, Jiahao
Liang, Tian
Song, Linfeng
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
contents Large language models (LLMs) can refine their responses based on feedback, enabling self-improvement through iterative training or test-time refinement. However, existing methods predominantly focus on refinement within the same reasoning format, which may lead to non-correcting behaviors. We propose CaP, a novel approach that uses external tools to refine chain-of-thought (CoT) responses generated by the same or other LLMs. CaP employs a two-stage training process: supervised fine-tuning followed by preference optimization with DPO variants. Our observations highlight the critical role of preference optimization in enabling effective refinement. Additionally, we compare several sampling strategies to leverage CoT and tools at inference time. Experimental results demonstrate CaP's potential for effective cross-reasoning refinement and efficient inference.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16871
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Teaching LLMs to Refine with Tools
Yu, Dian
Zhang, Yuheng
Xu, Jiahao
Liang, Tian
Song, Linfeng
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
Computation and Language
Large language models (LLMs) can refine their responses based on feedback, enabling self-improvement through iterative training or test-time refinement. However, existing methods predominantly focus on refinement within the same reasoning format, which may lead to non-correcting behaviors. We propose CaP, a novel approach that uses external tools to refine chain-of-thought (CoT) responses generated by the same or other LLMs. CaP employs a two-stage training process: supervised fine-tuning followed by preference optimization with DPO variants. Our observations highlight the critical role of preference optimization in enabling effective refinement. Additionally, we compare several sampling strategies to leverage CoT and tools at inference time. Experimental results demonstrate CaP's potential for effective cross-reasoning refinement and efficient inference.
title Teaching LLMs to Refine with Tools
topic Computation and Language
url https://arxiv.org/abs/2412.16871