Teaching LLMs to Refine with Tools
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912165960613888 |
|---|---|
| author | Yu, Dian Zhang, Yuheng Xu, Jiahao Liang, Tian Song, Linfeng Tu, Zhaopeng Mi, Haitao Yu, Dong |
| author_facet | Yu, Dian Zhang, Yuheng Xu, Jiahao Liang, Tian Song, Linfeng Tu, Zhaopeng Mi, Haitao Yu, Dong |
| contents | Large language models (LLMs) can refine their responses based on feedback, enabling self-improvement through iterative training or test-time refinement. However, existing methods predominantly focus on refinement within the same reasoning format, which may lead to non-correcting behaviors. We propose CaP, a novel approach that uses external tools to refine chain-of-thought (CoT) responses generated by the same or other LLMs. CaP employs a two-stage training process: supervised fine-tuning followed by preference optimization with DPO variants. Our observations highlight the critical role of preference optimization in enabling effective refinement. Additionally, we compare several sampling strategies to leverage CoT and tools at inference time. Experimental results demonstrate CaP's potential for effective cross-reasoning refinement and efficient inference. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_16871 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Teaching LLMs to Refine with Tools Yu, Dian Zhang, Yuheng Xu, Jiahao Liang, Tian Song, Linfeng Tu, Zhaopeng Mi, Haitao Yu, Dong Computation and Language Large language models (LLMs) can refine their responses based on feedback, enabling self-improvement through iterative training or test-time refinement. However, existing methods predominantly focus on refinement within the same reasoning format, which may lead to non-correcting behaviors. We propose CaP, a novel approach that uses external tools to refine chain-of-thought (CoT) responses generated by the same or other LLMs. CaP employs a two-stage training process: supervised fine-tuning followed by preference optimization with DPO variants. Our observations highlight the critical role of preference optimization in enabling effective refinement. Additionally, we compare several sampling strategies to leverage CoT and tools at inference time. Experimental results demonstrate CaP's potential for effective cross-reasoning refinement and efficient inference. |
| title | Teaching LLMs to Refine with Tools |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.16871 |