ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Doniparthi, Gajendra, Pandhare, Shashank Balu, Deßloch, Stefan, Mühlhaus, Timo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914205514334208
author Doniparthi, Gajendra
Pandhare, Shashank Balu
Deßloch, Stefan
Mühlhaus, Timo
author_facet Doniparthi, Gajendra
Pandhare, Shashank Balu
Deßloch, Stefan
Mühlhaus, Timo
contents Traditional search applications within Research Data Management (RDM) ecosystems are crucial in helping users discover and explore the structured metadata from the research datasets. Typically, text search engines require users to submit keyword-based queries rather than using natural language. However, using Large Language Models (LLMs) trained on domain-specific content for specialized natural language processing (NLP) tasks is becoming increasingly common. We present ArcBERT, an LLM-based system designed for integrated metadata exploration. ArcBERT understands natural language queries and relies on semantic matching, unlike traditional search applications. Notably, ArcBERT also understands the structure and hierarchies within the metadata, enabling it to handle diverse user querying patterns effectively.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata
Doniparthi, Gajendra
Pandhare, Shashank Balu
Deßloch, Stefan
Mühlhaus, Timo
Databases
Information Retrieval
Traditional search applications within Research Data Management (RDM) ecosystems are crucial in helping users discover and explore the structured metadata from the research datasets. Typically, text search engines require users to submit keyword-based queries rather than using natural language. However, using Large Language Models (LLMs) trained on domain-specific content for specialized natural language processing (NLP) tasks is becoming increasingly common. We present ArcBERT, an LLM-based system designed for integrated metadata exploration. ArcBERT understands natural language queries and relies on semantic matching, unlike traditional search applications. Notably, ArcBERT also understands the structure and hierarchies within the metadata, enabling it to handle diverse user querying patterns effectively.
title ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata
topic Databases
Information Retrieval
url https://arxiv.org/abs/2512.15365