Extracting infromation from text using an LLM model – first tests

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Carraro, Francesco
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901188324098048
author Carraro, Francesco
author_facet Carraro, Francesco
contents <p>This work investigates the use of <strong>Large Language Models </strong>(<strong>LLMs</strong>) to convert unstructured natural‑language instructions into structured JSON metadata suitable for scientific workflows. The system architecture combines a locally executed LLM via <strong>Ollama</strong>, a .NET Web API responsible for prompting and validation, and a lightweight console client. The processing pipeline operates in two stages: a linguistic normalization step that translates operator input into clear, unambiguous English, followed by schema‑guided extraction that enforces strict JSON structure. Through iterative prompt engineering, the approach achieves deterministic, schema‑compliant output while avoiding free text and hallucinated fields. Experiments show that smaller, instruction‑obedient models such as Phi‑3 provide the most reliable behavior under strong constraints. The resulting workflow is robust, offline‑capable, and well‑suited to institutional environments where reproducibility and metadata quality are essential. Future extensions may include schema expansion, ontology integration, and confidence scoring.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19184108
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Extracting infromation from text using an LLM model – first tests
Carraro, Francesco
LLM
JSON
Web API
Extract information
<p>This work investigates the use of <strong>Large Language Models </strong>(<strong>LLMs</strong>) to convert unstructured natural‑language instructions into structured JSON metadata suitable for scientific workflows. The system architecture combines a locally executed LLM via <strong>Ollama</strong>, a .NET Web API responsible for prompting and validation, and a lightweight console client. The processing pipeline operates in two stages: a linguistic normalization step that translates operator input into clear, unambiguous English, followed by schema‑guided extraction that enforces strict JSON structure. Through iterative prompt engineering, the approach achieves deterministic, schema‑compliant output while avoiding free text and hallucinated fields. Experiments show that smaller, instruction‑obedient models such as Phi‑3 provide the most reliable behavior under strong constraints. The resulting workflow is robust, offline‑capable, and well‑suited to institutional environments where reproducibility and metadata quality are essential. Future extensions may include schema expansion, ontology integration, and confidence scoring.</p>
title Extracting infromation from text using an LLM model – first tests
topic LLM
JSON
Web API
Extract information
url https://doi.org/10.5281/zenodo.19184108