INDUS: Effective and Efficient Language Models for Scientific Applications
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909372963094528 |
|---|---|
| author | Bhattacharjee, Bishwaranjan Trivedi, Aashka Muraoka, Masayasu Ramasubramanian, Muthukumaran Udagawa, Takuma Gurung, Iksha Pantha, Nishan Zhang, Rong Dandala, Bharath Ramachandran, Rahul Maskey, Manil Bugbee, Kaylin Little, Mike Fancher, Elizabeth Gerasimov, Irina Mehrabian, Armin Sanders, Lauren Costes, Sylvain Blanco-Cuaresma, Sergi Lockhart, Kelly Allen, Thomas Grezes, Felix Ansdell, Megan Accomazzi, Alberto El-Kurdi, Yousef Wertheimer, Davis Pfitzmann, Birgit Ramis, Cesar Berrospi Dolfi, Michele de Lima, Rafael Teixeira Vagenas, Panagiotis Mukkavilli, S. Karthik Staar, Peter Vahidinia, Sanaz McGranaghan, Ryan Lee, Tsendgar |
| author_facet | Bhattacharjee, Bishwaranjan Trivedi, Aashka Muraoka, Masayasu Ramasubramanian, Muthukumaran Udagawa, Takuma Gurung, Iksha Pantha, Nishan Zhang, Rong Dandala, Bharath Ramachandran, Rahul Maskey, Manil Bugbee, Kaylin Little, Mike Fancher, Elizabeth Gerasimov, Irina Mehrabian, Armin Sanders, Lauren Costes, Sylvain Blanco-Cuaresma, Sergi Lockhart, Kelly Allen, Thomas Grezes, Felix Ansdell, Megan Accomazzi, Alberto El-Kurdi, Yousef Wertheimer, Davis Pfitzmann, Birgit Ramis, Cesar Berrospi Dolfi, Michele de Lima, Rafael Teixeira Vagenas, Panagiotis Mukkavilli, S. Karthik Staar, Peter Vahidinia, Sanaz McGranaghan, Ryan Lee, Tsendgar |
| contents | Large language models (LLMs) trained on general domain corpora showed remarkable results on natural language processing (NLP) tasks. However, previous research demonstrated LLMs trained using domain-focused corpora perform better on specialized tasks. Inspired by this insight, we developed INDUS, a comprehensive suite of LLMs tailored for the closely-related domains of Earth science, biology, physics, heliophysics, planetary sciences and astrophysics, and trained using curated scientific corpora drawn from diverse data sources. The suite of models include: (1) an encoder model trained using domain-specific vocabulary and corpora to address NLP tasks, (2) a contrastive-learning based text embedding model trained using a diverse set of datasets to address information retrieval tasks and (3) smaller versions of these models created using knowledge distillation for applications which have latency or resource constraints. We also created three new scientific benchmark datasets, CLIMATE-CHANGE NER (entity-recognition), NASA-QA (extractive QA) and NASA-IR (IR) to accelerate research in these multi-disciplinary fields. We show that our models outperform both general-purpose (RoBERTa) and domain-specific (SCIBERT) encoders on these new tasks as well as existing tasks in the domains of interest. Furthermore, we demonstrate the use of these models in two industrial settings -- as a retrieval model for large-scale vector search applications and in automatic content tagging systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_10725 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | INDUS: Effective and Efficient Language Models for Scientific Applications Bhattacharjee, Bishwaranjan Trivedi, Aashka Muraoka, Masayasu Ramasubramanian, Muthukumaran Udagawa, Takuma Gurung, Iksha Pantha, Nishan Zhang, Rong Dandala, Bharath Ramachandran, Rahul Maskey, Manil Bugbee, Kaylin Little, Mike Fancher, Elizabeth Gerasimov, Irina Mehrabian, Armin Sanders, Lauren Costes, Sylvain Blanco-Cuaresma, Sergi Lockhart, Kelly Allen, Thomas Grezes, Felix Ansdell, Megan Accomazzi, Alberto El-Kurdi, Yousef Wertheimer, Davis Pfitzmann, Birgit Ramis, Cesar Berrospi Dolfi, Michele de Lima, Rafael Teixeira Vagenas, Panagiotis Mukkavilli, S. Karthik Staar, Peter Vahidinia, Sanaz McGranaghan, Ryan Lee, Tsendgar Computation and Language Information Retrieval Large language models (LLMs) trained on general domain corpora showed remarkable results on natural language processing (NLP) tasks. However, previous research demonstrated LLMs trained using domain-focused corpora perform better on specialized tasks. Inspired by this insight, we developed INDUS, a comprehensive suite of LLMs tailored for the closely-related domains of Earth science, biology, physics, heliophysics, planetary sciences and astrophysics, and trained using curated scientific corpora drawn from diverse data sources. The suite of models include: (1) an encoder model trained using domain-specific vocabulary and corpora to address NLP tasks, (2) a contrastive-learning based text embedding model trained using a diverse set of datasets to address information retrieval tasks and (3) smaller versions of these models created using knowledge distillation for applications which have latency or resource constraints. We also created three new scientific benchmark datasets, CLIMATE-CHANGE NER (entity-recognition), NASA-QA (extractive QA) and NASA-IR (IR) to accelerate research in these multi-disciplinary fields. We show that our models outperform both general-purpose (RoBERTa) and domain-specific (SCIBERT) encoders on these new tasks as well as existing tasks in the domains of interest. Furthermore, we demonstrate the use of these models in two industrial settings -- as a retrieval model for large-scale vector search applications and in automatic content tagging systems. |
| title | INDUS: Effective and Efficient Language Models for Scientific Applications |
| topic | Computation and Language Information Retrieval |
| url | https://arxiv.org/abs/2405.10725 |