Application Of Large Language Models For The Extraction Of Information From Particle Accelerator Technical Documentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Qing, Ischebeck, Rasmus, Sapinski, Maruisz, Grycner, Adam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914017326399488
author Dai, Qing
Ischebeck, Rasmus
Sapinski, Maruisz
Grycner, Adam
author_facet Dai, Qing
Ischebeck, Rasmus
Sapinski, Maruisz
Grycner, Adam
contents The large set of technical documentation of legacy accelerator systems, coupled with the retirement of experienced personnel, underscores the urgent need for efficient methods to preserve and transfer specialized knowledge. This paper explores the application of large language models (LLMs), to automate and enhance the extraction of information from particle accelerator technical documents. By exploiting LLMs, we aim to address the challenges of knowledge retention, enabling the retrieval of domain expertise embedded in legacy documentation. We present initial results of adapting LLMs to this specialized domain. Our evaluation demonstrates the effectiveness of LLMs in extracting, summarizing, and organizing knowledge, significantly reducing the risk of losing valuable insights as personnel retire. Furthermore, we discuss the limitations of current LLMs, such as interpretability and handling of rare domain-specific terms, and propose strategies for improvement. This work highlights the potential of LLMs to play a pivotal role in preserving institutional knowledge and ensuring continuity in highly specialized fields.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Application Of Large Language Models For The Extraction Of Information From Particle Accelerator Technical Documentation
Dai, Qing
Ischebeck, Rasmus
Sapinski, Maruisz
Grycner, Adam
Information Retrieval
Artificial Intelligence
Accelerator Physics
The large set of technical documentation of legacy accelerator systems, coupled with the retirement of experienced personnel, underscores the urgent need for efficient methods to preserve and transfer specialized knowledge. This paper explores the application of large language models (LLMs), to automate and enhance the extraction of information from particle accelerator technical documents. By exploiting LLMs, we aim to address the challenges of knowledge retention, enabling the retrieval of domain expertise embedded in legacy documentation. We present initial results of adapting LLMs to this specialized domain. Our evaluation demonstrates the effectiveness of LLMs in extracting, summarizing, and organizing knowledge, significantly reducing the risk of losing valuable insights as personnel retire. Furthermore, we discuss the limitations of current LLMs, such as interpretability and handling of rare domain-specific terms, and propose strategies for improvement. This work highlights the potential of LLMs to play a pivotal role in preserving institutional knowledge and ensuring continuity in highly specialized fields.
title Application Of Large Language Models For The Extraction Of Information From Particle Accelerator Technical Documentation
topic Information Retrieval
Artificial Intelligence
Accelerator Physics
url https://arxiv.org/abs/2509.02227