Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Chengyu, Han, Jinyi, Ying, Yizhou, Chen, Aili, He, Qianyu, Zhao, Haokun, Xia, Sirui, Guo, Haoran, Liang, Jiaqing, Chen, Zulong, Li, Liangyue, Xiao, Yanghua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910654433067008
author Du, Chengyu
Han, Jinyi
Ying, Yizhou
Chen, Aili
He, Qianyu
Zhao, Haokun
Xia, Sirui
Guo, Haoran
Liang, Jiaqing
Chen, Zulong
Li, Liangyue
Xiao, Yanghua
author_facet Du, Chengyu
Han, Jinyi
Ying, Yizhou
Chen, Aili
He, Qianyu
Zhao, Haokun
Xia, Sirui
Guo, Haoran
Liang, Jiaqing
Chen, Zulong
Li, Liangyue
Xiao, Yanghua
contents Recent advancements in large language models (LLMs) have demonstrated that progressive refinement, rather than providing a single answer, results in more accurate and thoughtful outputs. However, existing methods often rely heavily on supervision signals to evaluate previous responses, making it difficult to assess output quality in more open-ended scenarios effectively. Additionally, these methods are typically designed for specific tasks, which limits their generalization to new domains. To address these limitations, we propose Progressive Thought Refinement (PTR), a framework that enables LLMs to refine their responses progressively. PTR operates in two phases: (1) Thought data construction stage: We propose a weak and strong model collaborative selection strategy to build a high-quality progressive refinement dataset to ensure logical consistency from thought to answers, and the answers are gradually refined in each round. (2) Thought-Mask Fine-Tuning Phase: We design a training structure to mask the "thought" and adjust loss weights to encourage LLMs to refine prior thought, teaching them to implicitly understand "how to improve" rather than "what is correct." Experimental results show that PTR significantly enhances LLM performance across ten diverse tasks (avg. from 49.6% to 53.5%) without task-specific fine-tuning. Notably, in more open-ended tasks, LLMs also demonstrate substantial improvements in the quality of responses beyond mere accuracy, suggesting that PTR truly teaches LLMs to self-improve over time.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13413
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
Du, Chengyu
Han, Jinyi
Ying, Yizhou
Chen, Aili
He, Qianyu
Zhao, Haokun
Xia, Sirui
Guo, Haoran
Liang, Jiaqing
Chen, Zulong
Li, Liangyue
Xiao, Yanghua
Computation and Language
Artificial Intelligence
Recent advancements in large language models (LLMs) have demonstrated that progressive refinement, rather than providing a single answer, results in more accurate and thoughtful outputs. However, existing methods often rely heavily on supervision signals to evaluate previous responses, making it difficult to assess output quality in more open-ended scenarios effectively. Additionally, these methods are typically designed for specific tasks, which limits their generalization to new domains. To address these limitations, we propose Progressive Thought Refinement (PTR), a framework that enables LLMs to refine their responses progressively. PTR operates in two phases: (1) Thought data construction stage: We propose a weak and strong model collaborative selection strategy to build a high-quality progressive refinement dataset to ensure logical consistency from thought to answers, and the answers are gradually refined in each round. (2) Thought-Mask Fine-Tuning Phase: We design a training structure to mask the "thought" and adjust loss weights to encourage LLMs to refine prior thought, teaching them to implicitly understand "how to improve" rather than "what is correct." Experimental results show that PTR significantly enhances LLM performance across ten diverse tasks (avg. from 49.6% to 53.5%) without task-specific fine-tuning. Notably, in more open-ended tasks, LLMs also demonstrate substantial improvements in the quality of responses beyond mere accuracy, suggesting that PTR truly teaches LLMs to self-improve over time.
title Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.13413