Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Shaojie, Wang, Zhaobin, Zhuo, Chengxiang, Lu, Hui, Hu, Bo, Li, Zang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909057020854272
author Zhu, Shaojie
Wang, Zhaobin
Zhuo, Chengxiang
Lu, Hui
Hu, Bo
Li, Zang
author_facet Zhu, Shaojie
Wang, Zhaobin
Zhuo, Chengxiang
Lu, Hui
Hu, Bo
Li, Zang
contents CoT (Chain-of-Thought) is a way to solve reasoning problems for LLMs . Recently, many researches appear for improving the CoT capability of LLMs. In this work, we also proposed Olapa-MCoT, which is a LLMs based on llama2-13B PLM for finetuning and alignment learning. During the alignment training, we proposed the SimRRHF algorithm and Incorrect Data Relearning and mainly focused on optimizing the Chinese mathematical reasoning ability of Olapa-MCoT. The experiment achieved significant results, with the accuracy of Chinese mathematical reasoning up to 50%, 36% rise compared to llama2-13B. In addition, the accuracy of English reasoning ability also increased by nearly 4%.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17535
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
Zhu, Shaojie
Wang, Zhaobin
Zhuo, Chengxiang
Lu, Hui
Hu, Bo
Li, Zang
Artificial Intelligence
Computation and Language
Human-Computer Interaction
CoT (Chain-of-Thought) is a way to solve reasoning problems for LLMs . Recently, many researches appear for improving the CoT capability of LLMs. In this work, we also proposed Olapa-MCoT, which is a LLMs based on llama2-13B PLM for finetuning and alignment learning. During the alignment training, we proposed the SimRRHF algorithm and Incorrect Data Relearning and mainly focused on optimizing the Chinese mathematical reasoning ability of Olapa-MCoT. The experiment achieved significant results, with the accuracy of Chinese mathematical reasoning up to 50%, 36% rise compared to llama2-13B. In addition, the accuracy of English reasoning ability also increased by nearly 4%.
title Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2312.17535