G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2022
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917591389306880 |
|---|---|
| author | Wan, Zhongwei Yin, Yichun Zhang, Wei Shi, Jiaxin Shang, Lifeng Chen, Guangyong Jiang, Xin Liu, Qun |
| author_facet | Wan, Zhongwei Yin, Yichun Zhang, Wei Shi, Jiaxin Shang, Lifeng Chen, Guangyong Jiang, Xin Liu, Qun |
| contents | Recently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora. However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. (2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance. To alleviate this problem, we propose a new framework of General Memory Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge. Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM. We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP can achieve SOTA results on all tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2212_03613 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks Wan, Zhongwei Yin, Yichun Zhang, Wei Shi, Jiaxin Shang, Lifeng Chen, Guangyong Jiang, Xin Liu, Qun Computation and Language Recently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora. However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. (2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance. To alleviate this problem, we propose a new framework of General Memory Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge. Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM. We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP can achieve SOTA results on all tasks. |
| title | G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2212.03613 |