Towards Chapter-to-Chapter Context-Aware Literary Translation via Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Linghao, An, Li, Ma, Xuezhe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913427920781312
author Jin, Linghao
An, Li
Ma, Xuezhe
author_facet Jin, Linghao
An, Li
Ma, Xuezhe
contents Discourse phenomena in existing document-level translation datasets are sparse, which has been a fundamental obstacle in the development of context-aware machine translation models. Moreover, most existing document-level corpora and context-aware machine translation methods rely on an unrealistic assumption on sentence-level alignments. To mitigate these issues, we first curate a novel dataset of Chinese-English literature, which consists of 160 books with intricate discourse structures. Then, we propose a more pragmatic and challenging setting for context-aware translation, termed chapter-to-chapter (Ch2Ch) translation, and investigate the performance of commonly-used machine translation models under this setting. Furthermore, we introduce a potential approach of finetuning large language models (LLMs) within the domain of Ch2Ch literary translation, yielding impressive improvements over baselines. Through our comprehensive analysis, we unveil that literary translation under the Ch2Ch setting is challenging in nature, with respect to both model learning methods and translation decoding algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08978
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Chapter-to-Chapter Context-Aware Literary Translation via Large Language Models
Jin, Linghao
An, Li
Ma, Xuezhe
Computation and Language
Machine Learning
Discourse phenomena in existing document-level translation datasets are sparse, which has been a fundamental obstacle in the development of context-aware machine translation models. Moreover, most existing document-level corpora and context-aware machine translation methods rely on an unrealistic assumption on sentence-level alignments. To mitigate these issues, we first curate a novel dataset of Chinese-English literature, which consists of 160 books with intricate discourse structures. Then, we propose a more pragmatic and challenging setting for context-aware translation, termed chapter-to-chapter (Ch2Ch) translation, and investigate the performance of commonly-used machine translation models under this setting. Furthermore, we introduce a potential approach of finetuning large language models (LLMs) within the domain of Ch2Ch literary translation, yielding impressive improvements over baselines. Through our comprehensive analysis, we unveil that literary translation under the Ch2Ch setting is challenging in nature, with respect to both model learning methods and translation decoding algorithms.
title Towards Chapter-to-Chapter Context-Aware Literary Translation via Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.08978