Extracting infromation from text using an LLM model – first tests
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901188324098048 |
|---|---|
| author | Carraro, Francesco |
| author_facet | Carraro, Francesco |
| contents | <p>This work investigates the use of <strong>Large Language Models </strong>(<strong>LLMs</strong>) to convert unstructured natural‑language instructions into structured JSON metadata suitable for scientific workflows. The system architecture combines a locally executed LLM via <strong>Ollama</strong>, a .NET Web API responsible for prompting and validation, and a lightweight console client. The processing pipeline operates in two stages: a linguistic normalization step that translates operator input into clear, unambiguous English, followed by schema‑guided extraction that enforces strict JSON structure. Through iterative prompt engineering, the approach achieves deterministic, schema‑compliant output while avoiding free text and hallucinated fields. Experiments show that smaller, instruction‑obedient models such as Phi‑3 provide the most reliable behavior under strong constraints. The resulting workflow is robust, offline‑capable, and well‑suited to institutional environments where reproducibility and metadata quality are essential. Future extensions may include schema expansion, ontology integration, and confidence scoring.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19184108 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Extracting infromation from text using an LLM model – first tests Carraro, Francesco LLM JSON Web API Extract information <p>This work investigates the use of <strong>Large Language Models </strong>(<strong>LLMs</strong>) to convert unstructured natural‑language instructions into structured JSON metadata suitable for scientific workflows. The system architecture combines a locally executed LLM via <strong>Ollama</strong>, a .NET Web API responsible for prompting and validation, and a lightweight console client. The processing pipeline operates in two stages: a linguistic normalization step that translates operator input into clear, unambiguous English, followed by schema‑guided extraction that enforces strict JSON structure. Through iterative prompt engineering, the approach achieves deterministic, schema‑compliant output while avoiding free text and hallucinated fields. Experiments show that smaller, instruction‑obedient models such as Phi‑3 provide the most reliable behavior under strong constraints. The resulting workflow is robust, offline‑capable, and well‑suited to institutional environments where reproducibility and metadata quality are essential. Future extensions may include schema expansion, ontology integration, and confidence scoring.</p> |
| title | Extracting infromation from text using an LLM model – first tests |
| topic | LLM JSON Web API Extract information |
| url | https://doi.org/10.5281/zenodo.19184108 |