QueEn: A Large Language Model for Quechua-English Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junhao, Shu, Peng, Li, Yiwei, Zhao, Huaqin, Jiang, Hanqi, Pan, Yi, Zhou, Yifan, Liu, Zhengliang, Howe, Lewis C, Liu, Tianming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915051473993728
author Chen, Junhao
Shu, Peng
Li, Yiwei
Zhao, Huaqin
Jiang, Hanqi
Pan, Yi
Zhou, Yifan
Liu, Zhengliang
Howe, Lewis C
Liu, Tianming
author_facet Chen, Junhao
Shu, Peng
Li, Yiwei
Zhao, Huaqin
Jiang, Hanqi
Pan, Yi
Zhou, Yifan
Liu, Zhengliang
Howe, Lewis C
Liu, Tianming
contents Recent studies show that large language models (LLMs) are powerful tools for working with natural language, bringing advances in many areas of computational linguistics. However, these models face challenges when applied to low-resource languages due to limited training data and difficulty in understanding cultural nuances. In this paper, we propose QueEn, a novel approach for Quechua-English translation that combines Retrieval-Augmented Generation (RAG) with parameter-efficient fine-tuning techniques. Our method leverages external linguistic resources through RAG and uses Low-Rank Adaptation (LoRA) for efficient model adaptation. Experimental results show that our approach substantially exceeds baseline models, with a BLEU score of 17.6 compared to 1.5 for standard GPT models. The integration of RAG with fine-tuning allows our system to address the challenges of low-resource language translation while maintaining computational efficiency. This work contributes to the broader goal of preserving endangered languages through advanced language technologies.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QueEn: A Large Language Model for Quechua-English Translation
Chen, Junhao
Shu, Peng
Li, Yiwei
Zhao, Huaqin
Jiang, Hanqi
Pan, Yi
Zhou, Yifan
Liu, Zhengliang
Howe, Lewis C
Liu, Tianming
Computation and Language
Artificial Intelligence
Recent studies show that large language models (LLMs) are powerful tools for working with natural language, bringing advances in many areas of computational linguistics. However, these models face challenges when applied to low-resource languages due to limited training data and difficulty in understanding cultural nuances. In this paper, we propose QueEn, a novel approach for Quechua-English translation that combines Retrieval-Augmented Generation (RAG) with parameter-efficient fine-tuning techniques. Our method leverages external linguistic resources through RAG and uses Low-Rank Adaptation (LoRA) for efficient model adaptation. Experimental results show that our approach substantially exceeds baseline models, with a BLEU score of 17.6 compared to 1.5 for standard GPT models. The integration of RAG with fine-tuning allows our system to address the challenges of low-resource language translation while maintaining computational efficiency. This work contributes to the broader goal of preserving endangered languages through advanced language technologies.
title QueEn: A Large Language Model for Quechua-English Translation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.05184