Evaluation of LLM-based Strategies for the Extraction of Food Product Information from Online Shops

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brosch, Christoph, Brumm, Sian, Krieger, Rolf, Scheffler, Jonas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918084741169152
author Brosch, Christoph
Brumm, Sian
Krieger, Rolf
Scheffler, Jonas
author_facet Brosch, Christoph
Brumm, Sian
Krieger, Rolf
Scheffler, Jonas
contents Generative AI and large language models (LLMs) offer significant potential for automating the extraction of structured information from web pages. In this work, we focus on food product pages from online retailers and explore schema-constrained extraction approaches to retrieve key product attributes, such as ingredient lists and nutrition tables. We compare two LLM-based approaches, direct extraction and indirect extraction via generated functions, evaluating them in terms of accuracy, efficiency, and cost on a curated dataset of 3,000 food product pages from three different online shops. Our results show that although the indirect approach achieves slightly lower accuracy (96.48\%, $-1.61\%$ compared to direct extraction), it reduces the number of required LLM calls by 95.82\%, leading to substantial efficiency gains and lower operational costs. These findings suggest that indirect extraction approaches can provide scalable and cost-effective solutions for large-scale information extraction tasks from template-based web pages using LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation of LLM-based Strategies for the Extraction of Food Product Information from Online Shops
Brosch, Christoph
Brumm, Sian
Krieger, Rolf
Scheffler, Jonas
Computation and Language
Information Retrieval
Machine Learning
Generative AI and large language models (LLMs) offer significant potential for automating the extraction of structured information from web pages. In this work, we focus on food product pages from online retailers and explore schema-constrained extraction approaches to retrieve key product attributes, such as ingredient lists and nutrition tables. We compare two LLM-based approaches, direct extraction and indirect extraction via generated functions, evaluating them in terms of accuracy, efficiency, and cost on a curated dataset of 3,000 food product pages from three different online shops. Our results show that although the indirect approach achieves slightly lower accuracy (96.48\%, $-1.61\%$ compared to direct extraction), it reduces the number of required LLM calls by 95.82\%, leading to substantial efficiency gains and lower operational costs. These findings suggest that indirect extraction approaches can provide scalable and cost-effective solutions for large-scale information extraction tasks from template-based web pages using LLMs.
title Evaluation of LLM-based Strategies for the Extraction of Food Product Information from Online Shops
topic Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2506.21585