Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Hongyin, Peng, Hao, Lyu, Zhiheng, Hou, Lei, Li, Juanzi, Xiao, Jinghui
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929283234004992
author Zhu, Hongyin
Peng, Hao
Lyu, Zhiheng
Hou, Lei
Li, Juanzi
Xiao, Jinghui
author_facet Zhu, Hongyin
Peng, Hao
Lyu, Zhiheng
Hou, Lei
Li, Juanzi
Xiao, Jinghui
contents Existing technologies expand BERT from different perspectives, e.g. designing different pre-training tasks, different semantic granularities, and different model architectures. Few models consider expanding BERT from different text formats. In this paper, we propose a heterogeneous knowledge language model (\textbf{HKLM}), a unified pre-trained language model (PLM) for all forms of text, including unstructured text, semi-structured text, and well-structured text. To capture the corresponding relations among these multi-format knowledge, our approach uses masked language model objective to learn word knowledge, uses triple classification objective and title matching objective to learn entity knowledge and topic knowledge respectively. To obtain the aforementioned multi-format text, we construct a corpus in the tourism domain and conduct experiments on 5 tourism NLP datasets. The results show that our approach outperforms the pre-training of plain text using only 1/4 of the data. We further pre-train the domain-agnostic HKLM and achieve performance gains on the XNLI dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2109_01048
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation
Zhu, Hongyin
Peng, Hao
Lyu, Zhiheng
Hou, Lei
Li, Juanzi
Xiao, Jinghui
Computation and Language
Existing technologies expand BERT from different perspectives, e.g. designing different pre-training tasks, different semantic granularities, and different model architectures. Few models consider expanding BERT from different text formats. In this paper, we propose a heterogeneous knowledge language model (\textbf{HKLM}), a unified pre-trained language model (PLM) for all forms of text, including unstructured text, semi-structured text, and well-structured text. To capture the corresponding relations among these multi-format knowledge, our approach uses masked language model objective to learn word knowledge, uses triple classification objective and title matching objective to learn entity knowledge and topic knowledge respectively. To obtain the aforementioned multi-format text, we construct a corpus in the tourism domain and conduct experiments on 5 tourism NLP datasets. The results show that our approach outperforms the pre-training of plain text using only 1/4 of the data. We further pre-train the domain-agnostic HKLM and achieve performance gains on the XNLI dataset.
title Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation
topic Computation and Language
url https://arxiv.org/abs/2109.01048