Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Zicheng, Liang, Tian, Xu, Jiahao, Lin, Qiuzhi, Wang, Xing, Luo, Ruilin, Shi, Chufan, Li, Siheng, Yang, Yujiu, Tu, Zhaopeng
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910781738582016
author Lin, Zicheng
Liang, Tian
Xu, Jiahao
Lin, Qiuzhi
Wang, Xing
Luo, Ruilin
Shi, Chufan
Li, Siheng
Yang, Yujiu
Tu, Zhaopeng
author_facet Lin, Zicheng
Liang, Tian
Xu, Jiahao
Lin, Qiuzhi
Wang, Xing
Luo, Ruilin
Shi, Chufan
Li, Siheng
Yang, Yujiu
Tu, Zhaopeng
contents Mathematical reasoning tasks pose significant challenges for large language models (LLMs) because they require precise logical deduction and sequence analysis. In this work, we introduce the concept of critical tokens -- elements within reasoning trajectories that significantly influence incorrect outcomes. We present a novel framework for identifying these tokens through rollout sampling and demonstrate their substantial divergence from traditional error tokens. Through extensive experiments on datasets such as GSM8K and MATH500, we show that identifying and replacing critical tokens significantly improves model accuracy. We propose an efficient methodology for pinpointing these tokens in large-scale datasets using contrastive estimation and extend this framework to enhance model training processes with direct preference optimization (DPO). Experimental results on GSM8K and MATH500 benchmarks with the widely used models Llama-3 (8B and 70B) and Deepseek-math (7B) demonstrate the effectiveness of the proposed approach, cDPO. Our results underscore the potential of leveraging critical tokens to reduce errors in reasoning tasks, advancing the development of AI systems capable of robust logical deduction. Our code, annotated datasets, and trained models are available at https://github.com/chenzhiling9954/Critical-Tokens-Matter to support and encourage future research in this promising field.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19943
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
Lin, Zicheng
Liang, Tian
Xu, Jiahao
Lin, Qiuzhi
Wang, Xing
Luo, Ruilin
Shi, Chufan
Li, Siheng
Yang, Yujiu
Tu, Zhaopeng
Computation and Language
Artificial Intelligence
Machine Learning
Mathematical reasoning tasks pose significant challenges for large language models (LLMs) because they require precise logical deduction and sequence analysis. In this work, we introduce the concept of critical tokens -- elements within reasoning trajectories that significantly influence incorrect outcomes. We present a novel framework for identifying these tokens through rollout sampling and demonstrate their substantial divergence from traditional error tokens. Through extensive experiments on datasets such as GSM8K and MATH500, we show that identifying and replacing critical tokens significantly improves model accuracy. We propose an efficient methodology for pinpointing these tokens in large-scale datasets using contrastive estimation and extend this framework to enhance model training processes with direct preference optimization (DPO). Experimental results on GSM8K and MATH500 benchmarks with the widely used models Llama-3 (8B and 70B) and Deepseek-math (7B) demonstrate the effectiveness of the proposed approach, cDPO. Our results underscore the potential of leveraging critical tokens to reduce errors in reasoning tasks, advancing the development of AI systems capable of robust logical deduction. Our code, annotated datasets, and trained models are available at https://github.com/chenzhiling9954/Critical-Tokens-Matter to support and encourage future research in this promising field.
title Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.19943