_version_ 1866909372963094528
author Bhattacharjee, Bishwaranjan
Trivedi, Aashka
Muraoka, Masayasu
Ramasubramanian, Muthukumaran
Udagawa, Takuma
Gurung, Iksha
Pantha, Nishan
Zhang, Rong
Dandala, Bharath
Ramachandran, Rahul
Maskey, Manil
Bugbee, Kaylin
Little, Mike
Fancher, Elizabeth
Gerasimov, Irina
Mehrabian, Armin
Sanders, Lauren
Costes, Sylvain
Blanco-Cuaresma, Sergi
Lockhart, Kelly
Allen, Thomas
Grezes, Felix
Ansdell, Megan
Accomazzi, Alberto
El-Kurdi, Yousef
Wertheimer, Davis
Pfitzmann, Birgit
Ramis, Cesar Berrospi
Dolfi, Michele
de Lima, Rafael Teixeira
Vagenas, Panagiotis
Mukkavilli, S. Karthik
Staar, Peter
Vahidinia, Sanaz
McGranaghan, Ryan
Lee, Tsendgar
author_facet Bhattacharjee, Bishwaranjan
Trivedi, Aashka
Muraoka, Masayasu
Ramasubramanian, Muthukumaran
Udagawa, Takuma
Gurung, Iksha
Pantha, Nishan
Zhang, Rong
Dandala, Bharath
Ramachandran, Rahul
Maskey, Manil
Bugbee, Kaylin
Little, Mike
Fancher, Elizabeth
Gerasimov, Irina
Mehrabian, Armin
Sanders, Lauren
Costes, Sylvain
Blanco-Cuaresma, Sergi
Lockhart, Kelly
Allen, Thomas
Grezes, Felix
Ansdell, Megan
Accomazzi, Alberto
El-Kurdi, Yousef
Wertheimer, Davis
Pfitzmann, Birgit
Ramis, Cesar Berrospi
Dolfi, Michele
de Lima, Rafael Teixeira
Vagenas, Panagiotis
Mukkavilli, S. Karthik
Staar, Peter
Vahidinia, Sanaz
McGranaghan, Ryan
Lee, Tsendgar
contents Large language models (LLMs) trained on general domain corpora showed remarkable results on natural language processing (NLP) tasks. However, previous research demonstrated LLMs trained using domain-focused corpora perform better on specialized tasks. Inspired by this insight, we developed INDUS, a comprehensive suite of LLMs tailored for the closely-related domains of Earth science, biology, physics, heliophysics, planetary sciences and astrophysics, and trained using curated scientific corpora drawn from diverse data sources. The suite of models include: (1) an encoder model trained using domain-specific vocabulary and corpora to address NLP tasks, (2) a contrastive-learning based text embedding model trained using a diverse set of datasets to address information retrieval tasks and (3) smaller versions of these models created using knowledge distillation for applications which have latency or resource constraints. We also created three new scientific benchmark datasets, CLIMATE-CHANGE NER (entity-recognition), NASA-QA (extractive QA) and NASA-IR (IR) to accelerate research in these multi-disciplinary fields. We show that our models outperform both general-purpose (RoBERTa) and domain-specific (SCIBERT) encoders on these new tasks as well as existing tasks in the domains of interest. Furthermore, we demonstrate the use of these models in two industrial settings -- as a retrieval model for large-scale vector search applications and in automatic content tagging systems.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10725
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle INDUS: Effective and Efficient Language Models for Scientific Applications
Bhattacharjee, Bishwaranjan
Trivedi, Aashka
Muraoka, Masayasu
Ramasubramanian, Muthukumaran
Udagawa, Takuma
Gurung, Iksha
Pantha, Nishan
Zhang, Rong
Dandala, Bharath
Ramachandran, Rahul
Maskey, Manil
Bugbee, Kaylin
Little, Mike
Fancher, Elizabeth
Gerasimov, Irina
Mehrabian, Armin
Sanders, Lauren
Costes, Sylvain
Blanco-Cuaresma, Sergi
Lockhart, Kelly
Allen, Thomas
Grezes, Felix
Ansdell, Megan
Accomazzi, Alberto
El-Kurdi, Yousef
Wertheimer, Davis
Pfitzmann, Birgit
Ramis, Cesar Berrospi
Dolfi, Michele
de Lima, Rafael Teixeira
Vagenas, Panagiotis
Mukkavilli, S. Karthik
Staar, Peter
Vahidinia, Sanaz
McGranaghan, Ryan
Lee, Tsendgar
Computation and Language
Information Retrieval
Large language models (LLMs) trained on general domain corpora showed remarkable results on natural language processing (NLP) tasks. However, previous research demonstrated LLMs trained using domain-focused corpora perform better on specialized tasks. Inspired by this insight, we developed INDUS, a comprehensive suite of LLMs tailored for the closely-related domains of Earth science, biology, physics, heliophysics, planetary sciences and astrophysics, and trained using curated scientific corpora drawn from diverse data sources. The suite of models include: (1) an encoder model trained using domain-specific vocabulary and corpora to address NLP tasks, (2) a contrastive-learning based text embedding model trained using a diverse set of datasets to address information retrieval tasks and (3) smaller versions of these models created using knowledge distillation for applications which have latency or resource constraints. We also created three new scientific benchmark datasets, CLIMATE-CHANGE NER (entity-recognition), NASA-QA (extractive QA) and NASA-IR (IR) to accelerate research in these multi-disciplinary fields. We show that our models outperform both general-purpose (RoBERTa) and domain-specific (SCIBERT) encoders on these new tasks as well as existing tasks in the domains of interest. Furthermore, we demonstrate the use of these models in two industrial settings -- as a retrieval model for large-scale vector search applications and in automatic content tagging systems.
title INDUS: Effective and Efficient Language Models for Scientific Applications
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2405.10725