Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Anand, Avinash, Prasad, Kritarth, Kirtani, Chhavi, Nair, Ashwin R, Nema, Manvendra Kumar, Jaiswal, Raj, Shah, Rajiv Ratn
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915078339559424
author Anand, Avinash
Prasad, Kritarth
Kirtani, Chhavi
Nair, Ashwin R
Nema, Manvendra Kumar
Jaiswal, Raj
Shah, Rajiv Ratn
author_facet Anand, Avinash
Prasad, Kritarth
Kirtani, Chhavi
Nair, Ashwin R
Nema, Manvendra Kumar
Jaiswal, Raj
Shah, Rajiv Ratn
contents Large Language Models (LLMs) excel in linguistic tasks but struggle with mathematical reasoning, particularly in non English languages like Hindi. This research aims to enhance the mathematical reasoning skills of smaller, resource efficient open-source LLMs in both Hindi and English. We evaluate models like OpenHathi 7B, LLaMA-2 7B, WizardMath 7B, Mistral 7B, LLeMMa 7B, MAmmoTH 7B, Gemini Pro, and GPT-4 using zero-shot, few-shot chain-of-thought (CoT) methods, and supervised fine-tuning. Our approach incorporates curriculum learning, progressively training models on increasingly difficult problems, a novel Decomposition Strategy to simplify complex arithmetic operations, and a Structured Solution Design that divides solutions into phases. Our experiments result in notable performance enhancements. WizardMath 7B exceeds Gemini's accuracy on English datasets by +6% and matches Gemini's performance on Hindi datasets. Adopting a bilingual approach that combines English and Hindi samples achieves results comparable to individual language models, demonstrating the capability to learn mathematical reasoning in both languages. This research highlights the potential for improving mathematical reasoning in open-source LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18415
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
Anand, Avinash
Prasad, Kritarth
Kirtani, Chhavi
Nair, Ashwin R
Nema, Manvendra Kumar
Jaiswal, Raj
Shah, Rajiv Ratn
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) excel in linguistic tasks but struggle with mathematical reasoning, particularly in non English languages like Hindi. This research aims to enhance the mathematical reasoning skills of smaller, resource efficient open-source LLMs in both Hindi and English. We evaluate models like OpenHathi 7B, LLaMA-2 7B, WizardMath 7B, Mistral 7B, LLeMMa 7B, MAmmoTH 7B, Gemini Pro, and GPT-4 using zero-shot, few-shot chain-of-thought (CoT) methods, and supervised fine-tuning. Our approach incorporates curriculum learning, progressively training models on increasingly difficult problems, a novel Decomposition Strategy to simplify complex arithmetic operations, and a Structured Solution Design that divides solutions into phases. Our experiments result in notable performance enhancements. WizardMath 7B exceeds Gemini's accuracy on English datasets by +6% and matches Gemini's performance on Hindi datasets. Adopting a bilingual approach that combines English and Hindi samples achieves results comparable to individual language models, demonstrating the capability to learn mathematical reasoning in both languages. This research highlights the potential for improving mathematical reasoning in open-source LLMs.
title Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.18415