Method for Aggregating Unstructured Data Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lazebnyi, Vsevolod, Tereshkina, Natalia, Shabarina, Maria, Fedorov, Dmitriy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910142222565376
author Lazebnyi, Vsevolod
Tereshkina, Natalia
Shabarina, Maria
Fedorov, Dmitriy
author_facet Lazebnyi, Vsevolod
Tereshkina, Natalia
Shabarina, Maria
Fedorov, Dmitriy
contents This paper presents a method for the automated collection and aggregation of unstructured data from diverse web sources, utilizing Large Language Models (LLMs). The primary challenge with existing techniques is their instability when the structure of webpages changes, their limited support for dynamically loaded content during information collection, and the requirement for labor-intensive manual design of data pre-processing processes. The proposed algorithm integrates hybrid web scraping (Goose3 for static pages and Selenium+WebDriver for dynamic ones), data storage in a non-relational MongoDB database management system (DBMS), and intelligent extraction and normalization of information using LLMs into a predetermined JSON schema. A key scientific contribution of this study is a two-stage verification process for the generated data, designed to eliminate potential hallucinations byy comparing the embeddings of multiple LLM outputs obtained with different temperature parameter values, combined with formalized rules for monitoring data consistency and integrity. The experimental findings indicate a high level of accuracy in the completion of key fields, as well as the robustness of the proposed methodology to changes in web page structures. This makes it suitable for use in tasks such as news content aggregation, monitoring, and log analysis in near real-time mode, with the capacity to scale rapidly in terms of the number of sources.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16425
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Method for Aggregating Unstructured Data Using Large Language Models
Lazebnyi, Vsevolod
Tereshkina, Natalia
Shabarina, Maria
Fedorov, Dmitriy
Databases
Machine Learning
68T50
I.2.7; H.3.3; H.2.8
This paper presents a method for the automated collection and aggregation of unstructured data from diverse web sources, utilizing Large Language Models (LLMs). The primary challenge with existing techniques is their instability when the structure of webpages changes, their limited support for dynamically loaded content during information collection, and the requirement for labor-intensive manual design of data pre-processing processes. The proposed algorithm integrates hybrid web scraping (Goose3 for static pages and Selenium+WebDriver for dynamic ones), data storage in a non-relational MongoDB database management system (DBMS), and intelligent extraction and normalization of information using LLMs into a predetermined JSON schema. A key scientific contribution of this study is a two-stage verification process for the generated data, designed to eliminate potential hallucinations byy comparing the embeddings of multiple LLM outputs obtained with different temperature parameter values, combined with formalized rules for monitoring data consistency and integrity. The experimental findings indicate a high level of accuracy in the completion of key fields, as well as the robustness of the proposed methodology to changes in web page structures. This makes it suitable for use in tasks such as news content aggregation, monitoring, and log analysis in near real-time mode, with the capacity to scale rapidly in terms of the number of sources.
title Method for Aggregating Unstructured Data Using Large Language Models
topic Databases
Machine Learning
68T50
I.2.7; H.3.3; H.2.8
url https://arxiv.org/abs/2604.16425