Using Language Models to Disambiguate Lexical Choices in Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912111842557952 |
|---|---|
| author | Barua, Josh Subramanian, Sanjay Yin, Kayo Suhr, Alane |
| author_facet | Barua, Josh Subramanian, Sanjay Yin, Kayo Suhr, Alane |
| contents | In translation, a concept represented by a single word in a source language can have multiple variations in a target language. The task of lexical selection requires using context to identify which variation is most appropriate for a source text. We work with native speakers of nine languages to create DTAiLS, a dataset of 1,377 sentence pairs that exhibit cross-lingual concept variation when translating from English. We evaluate recent LLMs and neural machine translation systems on DTAiLS, with the best-performing model, GPT-4, achieving from 67 to 85% accuracy across languages. Finally, we use language models to generate English rules describing target-language concept variations. Providing weaker models with high-quality lexical rules improves accuracy substantially, in some cases reaching or outperforming GPT-4. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_05781 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Using Language Models to Disambiguate Lexical Choices in Translation Barua, Josh Subramanian, Sanjay Yin, Kayo Suhr, Alane Computation and Language Artificial Intelligence In translation, a concept represented by a single word in a source language can have multiple variations in a target language. The task of lexical selection requires using context to identify which variation is most appropriate for a source text. We work with native speakers of nine languages to create DTAiLS, a dataset of 1,377 sentence pairs that exhibit cross-lingual concept variation when translating from English. We evaluate recent LLMs and neural machine translation systems on DTAiLS, with the best-performing model, GPT-4, achieving from 67 to 85% accuracy across languages. Finally, we use language models to generate English rules describing target-language concept variations. Providing weaker models with high-quality lexical rules improves accuracy substantially, in some cases reaching or outperforming GPT-4. |
| title | Using Language Models to Disambiguate Lexical Choices in Translation |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2411.05781 |