VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Phan, Hoang Hai, Vu, Nguyen Duc Minh, Phuong, Nam Dang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911204976361472
author Phan, Hoang Hai
Vu, Nguyen Duc Minh
Phuong, Nam Dang
author_facet Phan, Hoang Hai
Vu, Nguyen Duc Minh
Phuong, Nam Dang
contents Neural Machine Translation (NMT) driven by Transformer architectures has advanced significantly, yet faces challenges with low-resource language pairs like Vietnamese-Japanese (Vi-Ja). Issues include sparse parallel data and handling linguistic/cultural nuances. Recent progress in Large Language Models (LLMs) with strong reasoning, often refined via Reinforcement Learning (RL), enables high-quality synthetic data generation. We introduce VNJPTranslate, a pipeline designed to systematically address the Vi-Ja translation task. It features a targeted data augmentation strategy using advanced LLMs with Chain-of-Thought prompting for challenging segments identified via corpus analysis. Subsequently, we employ efficient fine-tuning techniques (Unsloth with QLoRA) on a capable, low-parameter autoregressive model (specifically, a fine-tuned version of the 1.8B parameter Sailor model, which is based on the Qwen architecture) to create a practical and high-performing translation system. This integrated approach aims to improve Vi-Ja translation quality significantly over existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation
Phan, Hoang Hai
Vu, Nguyen Duc Minh
Phuong, Nam Dang
Computation and Language
Artificial Intelligence
Neural Machine Translation (NMT) driven by Transformer architectures has advanced significantly, yet faces challenges with low-resource language pairs like Vietnamese-Japanese (Vi-Ja). Issues include sparse parallel data and handling linguistic/cultural nuances. Recent progress in Large Language Models (LLMs) with strong reasoning, often refined via Reinforcement Learning (RL), enables high-quality synthetic data generation. We introduce VNJPTranslate, a pipeline designed to systematically address the Vi-Ja translation task. It features a targeted data augmentation strategy using advanced LLMs with Chain-of-Thought prompting for challenging segments identified via corpus analysis. Subsequently, we employ efficient fine-tuning techniques (Unsloth with QLoRA) on a capable, low-parameter autoregressive model (specifically, a fine-tuned version of the 1.8B parameter Sailor model, which is based on the Qwen architecture) to create a practical and high-performing translation system. This integrated approach aims to improve Vi-Ja translation quality significantly over existing baselines.
title VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.00339