SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: She, Fengyu, Wang, Nan, Wu, Hongfei, Wan, Ziyi, Wang, Jingmian, Wang, Chang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908528468295680
author She, Fengyu
Wang, Nan
Wu, Hongfei
Wan, Ziyi
Wang, Jingmian
Wang, Chang
author_facet She, Fengyu
Wang, Nan
Wu, Hongfei
Wan, Ziyi
Wang, Jingmian
Wang, Chang
contents Scientific literature is growing exponentially, creating a critical bottleneck for researchers to efficiently synthesize knowledge. While general-purpose Large Language Models (LLMs) show potential in text processing, they often fail to capture scientific domain-specific nuances (e.g., technical jargon, methodological rigor) and struggle with complex scientific tasks, limiting their utility for interdisciplinary research. To address these gaps, this paper presents SciGPT, a domain-adapted foundation model for scientific literature understanding and ScienceBench, an open source benchmark tailored to evaluate scientific LLMs. Built on the Qwen3 architecture, SciGPT incorporates three key innovations: (1) low-cost domain distillation via a two-stage pipeline to balance performance and efficiency; (2) a Sparse Mixture-of-Experts (SMoE) attention mechanism that cuts memory consumption by 55\% for 32,000-token long-document reasoning; and (3) knowledge-aware adaptation integrating domain ontologies to bridge interdisciplinary knowledge gaps. Experimental results on ScienceBench show that SciGPT outperforms GPT-4o in core scientific tasks including sequence labeling, generation, and inference. It also exhibits strong robustness in unseen scientific tasks, validating its potential to facilitate AI-augmented scientific discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery
She, Fengyu
Wang, Nan
Wu, Hongfei
Wan, Ziyi
Wang, Jingmian
Wang, Chang
Computation and Language
Scientific literature is growing exponentially, creating a critical bottleneck for researchers to efficiently synthesize knowledge. While general-purpose Large Language Models (LLMs) show potential in text processing, they often fail to capture scientific domain-specific nuances (e.g., technical jargon, methodological rigor) and struggle with complex scientific tasks, limiting their utility for interdisciplinary research. To address these gaps, this paper presents SciGPT, a domain-adapted foundation model for scientific literature understanding and ScienceBench, an open source benchmark tailored to evaluate scientific LLMs. Built on the Qwen3 architecture, SciGPT incorporates three key innovations: (1) low-cost domain distillation via a two-stage pipeline to balance performance and efficiency; (2) a Sparse Mixture-of-Experts (SMoE) attention mechanism that cuts memory consumption by 55\% for 32,000-token long-document reasoning; and (3) knowledge-aware adaptation integrating domain ontologies to bridge interdisciplinary knowledge gaps. Experimental results on ScienceBench show that SciGPT outperforms GPT-4o in core scientific tasks including sequence labeling, generation, and inference. It also exhibits strong robustness in unseen scientific tasks, validating its potential to facilitate AI-augmented scientific discovery.
title SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery
topic Computation and Language
url https://arxiv.org/abs/2509.08032