GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tung, Luu Quy, Viet, Hoang Quoc, Loc, Pham Bao, Thu, Vo Trong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908483274670080
author Tung, Luu Quy
Viet, Hoang Quoc
Loc, Pham Bao
Thu, Vo Trong
author_facet Tung, Luu Quy
Viet, Hoang Quoc
Loc, Pham Bao
Thu, Vo Trong
contents Chain-of-Thought (CoT) is a robust approach for tackling LLM tasks that require intermediate reasoning steps prior to generating a final answer. In this paper, we present GreenMind-Medium-14B-R1, the Vietnamese reasoning model inspired by the finetuning strategy based on Group Relative Policy Optimization. We also leverage a high-quality Vietnamese synthesized reasoning dataset and design two reward functions to tackle the main limitations of this technique: (i) language mixing, where we explicitly detect the presence of biased language characters during the process of sampling tokens, and (ii) we leverage Sentence Transformer-based models to ensure that the generated reasoning content maintains factual correctness and does not distort the final output. Experimental results on the Vietnamese dataset from the VLSP 2023 Challenge demonstrate that our model outperforms prior works and enhances linguistic consistency in its responses. Furthermore, we extend our evaluation to SeaExam-a multilingual multiple-choice dataset, showing the effectiveness of our reasoning method compared to few-shot prompting techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning
Tung, Luu Quy
Viet, Hoang Quoc
Loc, Pham Bao
Thu, Vo Trong
Computation and Language
Chain-of-Thought (CoT) is a robust approach for tackling LLM tasks that require intermediate reasoning steps prior to generating a final answer. In this paper, we present GreenMind-Medium-14B-R1, the Vietnamese reasoning model inspired by the finetuning strategy based on Group Relative Policy Optimization. We also leverage a high-quality Vietnamese synthesized reasoning dataset and design two reward functions to tackle the main limitations of this technique: (i) language mixing, where we explicitly detect the presence of biased language characters during the process of sampling tokens, and (ii) we leverage Sentence Transformer-based models to ensure that the generated reasoning content maintains factual correctness and does not distort the final output. Experimental results on the Vietnamese dataset from the VLSP 2023 Challenge demonstrate that our model outperforms prior works and enhances linguistic consistency in its responses. Furthermore, we extend our evaluation to SeaExam-a multilingual multiple-choice dataset, showing the effectiveness of our reasoning method compared to few-shot prompting techniques.
title GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning
topic Computation and Language
url https://arxiv.org/abs/2504.16832