Enhancing Large Language Models with Domain-Specific Knowledge: The Case in Topological Materials

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, HuangChao, Zhang, Baohua, Jin, Zhong, Zhu, Tiannian, Wu, Quansheng, Weng, Hongming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917137622237184
author Xu, HuangChao
Zhang, Baohua
Jin, Zhong
Zhu, Tiannian
Wu, Quansheng
Weng, Hongming
author_facet Xu, HuangChao
Zhang, Baohua
Jin, Zhong
Zhu, Tiannian
Wu, Quansheng
Weng, Hongming
contents Large language models (LLMs), such as ChatGPT, have demonstrated impressive performance in the text generation task, showing the ability to understand and respond to complex instructions. However, the performance of naive LLMs in speciffc domains is limited due to the scarcity of domain-speciffc corpora and specialized training. Moreover, training a specialized large-scale model necessitates signiffcant hardware resources, which restricts researchers from leveraging such models to drive advances. Hence, it is crucial to further improve and optimize LLMs to meet speciffc domain demands and enhance their scalability. Based on the condensed matter data center, we establish a material knowledge graph (MaterialsKG) and integrate it with literature. Using large language models and prompt learning, we develop a specialized dialogue system for topological materials called TopoChat. Compared to naive LLMs, TopoChat exhibits superior performance in structural and property querying, material recommendation, and complex relational reasoning. This system enables efffcient and precise retrieval of information and facilitates knowledge interaction, thereby encouraging the advancement on the ffeld of condensed matter materials.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13732
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Large Language Models with Domain-Specific Knowledge: The Case in Topological Materials
Xu, HuangChao
Zhang, Baohua
Jin, Zhong
Zhu, Tiannian
Wu, Quansheng
Weng, Hongming
Computation and Language
Materials Science
Machine Learning
Large language models (LLMs), such as ChatGPT, have demonstrated impressive performance in the text generation task, showing the ability to understand and respond to complex instructions. However, the performance of naive LLMs in speciffc domains is limited due to the scarcity of domain-speciffc corpora and specialized training. Moreover, training a specialized large-scale model necessitates signiffcant hardware resources, which restricts researchers from leveraging such models to drive advances. Hence, it is crucial to further improve and optimize LLMs to meet speciffc domain demands and enhance their scalability. Based on the condensed matter data center, we establish a material knowledge graph (MaterialsKG) and integrate it with literature. Using large language models and prompt learning, we develop a specialized dialogue system for topological materials called TopoChat. Compared to naive LLMs, TopoChat exhibits superior performance in structural and property querying, material recommendation, and complex relational reasoning. This system enables efffcient and precise retrieval of information and facilitates knowledge interaction, thereby encouraging the advancement on the ffeld of condensed matter materials.
title Enhancing Large Language Models with Domain-Specific Knowledge: The Case in Topological Materials
topic Computation and Language
Materials Science
Machine Learning
url https://arxiv.org/abs/2409.13732