JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nagata, Masaaki, Chousa, Katsuki, Yasuda, Norihito
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909749026488320
author Nagata, Masaaki
Chousa, Katsuki
Yasuda, Norihito
author_facet Nagata, Masaaki
Chousa, Katsuki
Yasuda, Norihito
contents We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021. We obtained the publication of unexamined patent applications from the Japan Patent Office (JPO) and the United States Patent and Trademark Office (USPTO). We also obtained patent family information from the DOCDB, that is a bibliographic database maintained by the European Patent Office (EPO). We extracted approximately 1.4M Japanese-English document pairs, which are translations of each other based on the patent families, and extracted about 350M sentence pairs from the document pairs using a translation-based sentence alignment method whose initial translation model is bootstrapped from a dictionary-based sentence alignment method. We experimentally improved the accuracy of the patent translations by 20 bleu points by adding more than 300M sentence pairs obtained from patent applications to 22M sentence pairs obtained from the web.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16303
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
Nagata, Masaaki
Chousa, Katsuki
Yasuda, Norihito
Computation and Language
We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021. We obtained the publication of unexamined patent applications from the Japan Patent Office (JPO) and the United States Patent and Trademark Office (USPTO). We also obtained patent family information from the DOCDB, that is a bibliographic database maintained by the European Patent Office (EPO). We extracted approximately 1.4M Japanese-English document pairs, which are translations of each other based on the patent families, and extracted about 350M sentence pairs from the document pairs using a translation-based sentence alignment method whose initial translation model is bootstrapped from a dictionary-based sentence alignment method. We experimentally improved the accuracy of the patent translations by 20 bleu points by adding more than 300M sentence pairs obtained from patent applications to 22M sentence pairs obtained from the web.
title JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
topic Computation and Language
url https://arxiv.org/abs/2508.16303