Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field
Fuente:
arXiv
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866929439482314752 |
|---|---|
| author | Kerner, Tobias |
| author_facet | Kerner, Tobias |
| contents | There are many cases where LLMs are used for specific tasks in a single domain. These usually require less general, but more domain-specific knowledge. Highly capable, general-purpose state-of-the-art language models like GPT-4 or Claude-3-opus can often be used for such tasks, but they are very large and cannot be run locally, even if they were not proprietary. This can be a problem when working with sensitive data. This paper focuses on domain-specific and mixed-domain pretraining as potentially more efficient methods than general pretraining for specialized language models. We will take a look at work related to domain-specific pretraining, specifically in the medical area, and compare benchmark results of specialized language models to general-purpose language models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_14076 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field Kerner, Tobias Machine Learning Artificial Intelligence Computation and Language I.2.6; I.2.7 There are many cases where LLMs are used for specific tasks in a single domain. These usually require less general, but more domain-specific knowledge. Highly capable, general-purpose state-of-the-art language models like GPT-4 or Claude-3-opus can often be used for such tasks, but they are very large and cannot be run locally, even if they were not proprietary. This can be a problem when working with sensitive data. This paper focuses on domain-specific and mixed-domain pretraining as potentially more efficient methods than general pretraining for specialized language models. We will take a look at work related to domain-specific pretraining, specifically in the medical area, and compare benchmark results of specialized language models to general-purpose language models. |
| title | Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field |
| topic | Machine Learning Artificial Intelligence Computation and Language I.2.6; I.2.7 |
| url | https://arxiv.org/abs/2407.14076 |