Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Kaishuai, Yu, Tiezheng, Hou, Wenjun, Cheng, Yi, Leong, Chak Tou, Li, Liangyou, Jiang, Xin, Shang, Lifeng, Liu, Qun, Li, Wenjie
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915305815539712
author Xu, Kaishuai
Yu, Tiezheng
Hou, Wenjun
Cheng, Yi
Leong, Chak Tou
Li, Liangyou
Jiang, Xin
Shang, Lifeng
Liu, Qun
Li, Wenjie
author_facet Xu, Kaishuai
Yu, Tiezheng
Hou, Wenjun
Cheng, Yi
Leong, Chak Tou
Li, Liangyou
Jiang, Xin
Shang, Lifeng
Liu, Qun
Li, Wenjie
contents Large Language Models (LLMs) have exhibited strong mathematical reasoning prowess, tackling tasks ranging from basic arithmetic to advanced competition-level problems. However, frequently occurring subtle yet critical errors, such as miscalculations or incorrect substitutions, limit the LLMs' full potential. Existing studies to improve mathematical ability typically involve applying preference learning to step-wise solution pairs. Although these methods leverage samples of varying granularity to mitigate reasoning errors, they overlook critical subtle errors. In this work, we propose a novel preference learning framework called eRror-Injected Self-Editing (RISE), which injects predefined subtle errors into pivotal tokens in reasoning or computation steps to construct hard pairs for error mitigation. In detail, RISE uses the LLM itself to edit a small number of tokens in the solution, injecting designed subtle errors. Then, pairs composed of self-edited solutions and their corresponding correct ones, along with pairs of correct and incorrect solutions obtained through sampling, are used together for subtle error-aware DPO training. Compared with other preference learning methods, RISE further refines the training objective without requiring fine-grained sampling or preference annotation. Extensive experiments validate the effectiveness of RISE, with preference learning on Qwen2-7B-Instruct yielding notable improvements of 3.0% on GSM8K and 7.9% on MATH with only 4.5K training samples. Moreover, the effect of error mitigation extends from mathematical reasoning to logical reasoning and code generation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06638
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
Xu, Kaishuai
Yu, Tiezheng
Hou, Wenjun
Cheng, Yi
Leong, Chak Tou
Li, Liangyou
Jiang, Xin
Shang, Lifeng
Liu, Qun
Li, Wenjie
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have exhibited strong mathematical reasoning prowess, tackling tasks ranging from basic arithmetic to advanced competition-level problems. However, frequently occurring subtle yet critical errors, such as miscalculations or incorrect substitutions, limit the LLMs' full potential. Existing studies to improve mathematical ability typically involve applying preference learning to step-wise solution pairs. Although these methods leverage samples of varying granularity to mitigate reasoning errors, they overlook critical subtle errors. In this work, we propose a novel preference learning framework called eRror-Injected Self-Editing (RISE), which injects predefined subtle errors into pivotal tokens in reasoning or computation steps to construct hard pairs for error mitigation. In detail, RISE uses the LLM itself to edit a small number of tokens in the solution, injecting designed subtle errors. Then, pairs composed of self-edited solutions and their corresponding correct ones, along with pairs of correct and incorrect solutions obtained through sampling, are used together for subtle error-aware DPO training. Compared with other preference learning methods, RISE further refines the training objective without requiring fine-grained sampling or preference annotation. Extensive experiments validate the effectiveness of RISE, with preference learning on Qwen2-7B-Instruct yielding notable improvements of 3.0% on GSM8K and 7.9% on MATH with only 4.5K training samples. Moreover, the effect of error mitigation extends from mathematical reasoning to logical reasoning and code generation.
title Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.06638