Saved in:
Bibliographic Details
Main Authors: Fu, Yuxin, Si, Shijing, Mai, Leyi, Li, Xi-ang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.18856
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910503797784576
author Fu, Yuxin
Si, Shijing
Mai, Leyi
Li, Xi-ang
author_facet Fu, Yuxin
Si, Shijing
Mai, Leyi
Li, Xi-ang
contents Large Language Models (LLMs) have stunningly advanced the field of machine translation, though their effectiveness within the financial domain remains largely underexplored. To probe this issue, we constructed a fine-grained Chinese-English parallel corpus of financial news called FFN. We acquired financial news articles spanning between January 1st, 2014, to December 31, 2023, from mainstream media websites such as CNN, FOX, and China Daily. The dataset consists of 1,013 main text and 809 titles, all of which have been manually corrected. We measured the translation quality of two LLMs -- ChatGPT and ERNIE-bot, utilizing BLEU, TER and chrF scores as the evaluation metrics. For comparison, we also trained an OpenNMT model based on our dataset. We detail problems of LLMs and provide in-depth analysis, intending to stimulate further research and solutions in this largely uncharted territory. Our research underlines the need to optimize LLMs within the specific field of financial translation to ensure accuracy and quality.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18856
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FFN: a Fine-grained Chinese-English Financial Domain Parallel Corpus
Fu, Yuxin
Si, Shijing
Mai, Leyi
Li, Xi-ang
Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Large Language Models (LLMs) have stunningly advanced the field of machine translation, though their effectiveness within the financial domain remains largely underexplored. To probe this issue, we constructed a fine-grained Chinese-English parallel corpus of financial news called FFN. We acquired financial news articles spanning between January 1st, 2014, to December 31, 2023, from mainstream media websites such as CNN, FOX, and China Daily. The dataset consists of 1,013 main text and 809 titles, all of which have been manually corrected. We measured the translation quality of two LLMs -- ChatGPT and ERNIE-bot, utilizing BLEU, TER and chrF scores as the evaluation metrics. For comparison, we also trained an OpenNMT model based on our dataset. We detail problems of LLMs and provide in-depth analysis, intending to stimulate further research and solutions in this largely uncharted territory. Our research underlines the need to optimize LLMs within the specific field of financial translation to ensure accuracy and quality.
title FFN: a Fine-grained Chinese-English Financial Domain Parallel Corpus
topic Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2406.18856