Exploring the Impact of Corpus Diversity on Financial Pretrained Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Choe, Jaeyoung, Noh, Keonwoong, Kim, Nayeon, Ahn, Seyun, Jung, Woohwan
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913858481815552
author Choe, Jaeyoung
Noh, Keonwoong
Kim, Nayeon
Ahn, Seyun
Jung, Woohwan
author_facet Choe, Jaeyoung
Noh, Keonwoong
Kim, Nayeon
Ahn, Seyun
Jung, Woohwan
contents Over the past few years, various domain-specific pretrained language models (PLMs) have been proposed and have outperformed general-domain PLMs in specialized areas such as biomedical, scientific, and clinical domains. In addition, financial PLMs have been studied because of the high economic impact of financial data analysis. However, we found that financial PLMs were not pretrained on sufficiently diverse financial data. This lack of diverse training data leads to a subpar generalization performance, resulting in general-purpose PLMs, including BERT, often outperforming financial PLMs on many downstream tasks. To address this issue, we collected a broad range of financial corpus and trained the Financial Language Model (FiLM) on these diverse datasets. Our experimental results confirm that FiLM outperforms not only existing financial PLMs but also general domain PLMs. Furthermore, we provide empirical evidence that this improvement can be achieved even for unseen corpus groups.
format Preprint
id arxiv_https___arxiv_org_abs_2310_13312
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Exploring the Impact of Corpus Diversity on Financial Pretrained Language Models
Choe, Jaeyoung
Noh, Keonwoong
Kim, Nayeon
Ahn, Seyun
Jung, Woohwan
Computation and Language
Over the past few years, various domain-specific pretrained language models (PLMs) have been proposed and have outperformed general-domain PLMs in specialized areas such as biomedical, scientific, and clinical domains. In addition, financial PLMs have been studied because of the high economic impact of financial data analysis. However, we found that financial PLMs were not pretrained on sufficiently diverse financial data. This lack of diverse training data leads to a subpar generalization performance, resulting in general-purpose PLMs, including BERT, often outperforming financial PLMs on many downstream tasks. To address this issue, we collected a broad range of financial corpus and trained the Financial Language Model (FiLM) on these diverse datasets. Our experimental results confirm that FiLM outperforms not only existing financial PLMs but also general domain PLMs. Furthermore, we provide empirical evidence that this improvement can be achieved even for unseen corpus groups.
title Exploring the Impact of Corpus Diversity on Financial Pretrained Language Models
topic Computation and Language
url https://arxiv.org/abs/2310.13312