Building a Macedonian Recipe Dataset: Collection, Parsing, and Comparative Analysis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sasanski, Darko, Peshevski, Dimitar, Stojanov, Riste, Trajanov, Dimitar
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911269493145600
author Sasanski, Darko
Peshevski, Dimitar
Stojanov, Riste
Trajanov, Dimitar
author_facet Sasanski, Darko
Peshevski, Dimitar
Stojanov, Riste
Trajanov, Dimitar
contents Computational gastronomy increasingly relies on diverse, high-quality recipe datasets to capture regional culinary traditions. Although there are large-scale collections for major languages, Macedonian recipes remain under-represented in digital research. In this work, we present the first systematic effort to construct a Macedonian recipe dataset through web scraping and structured parsing. We address challenges in processing heterogeneous ingredient descriptions, including unit, quantity, and descriptor normalization. An exploratory analysis of ingredient frequency and co-occurrence patterns, using measures such as Pointwise Mutual Information and Lift score, highlights distinctive ingredient combinations that characterize Macedonian cuisine. The resulting dataset contributes a new resource for studying food culture in underrepresented languages and offers insights into the unique patterns of Macedonian culinary tradition.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14128
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Building a Macedonian Recipe Dataset: Collection, Parsing, and Comparative Analysis
Sasanski, Darko
Peshevski, Dimitar
Stojanov, Riste
Trajanov, Dimitar
Computation and Language
Computational gastronomy increasingly relies on diverse, high-quality recipe datasets to capture regional culinary traditions. Although there are large-scale collections for major languages, Macedonian recipes remain under-represented in digital research. In this work, we present the first systematic effort to construct a Macedonian recipe dataset through web scraping and structured parsing. We address challenges in processing heterogeneous ingredient descriptions, including unit, quantity, and descriptor normalization. An exploratory analysis of ingredient frequency and co-occurrence patterns, using measures such as Pointwise Mutual Information and Lift score, highlights distinctive ingredient combinations that characterize Macedonian cuisine. The resulting dataset contributes a new resource for studying food culture in underrepresented languages and offers insights into the unique patterns of Macedonian culinary tradition.
title Building a Macedonian Recipe Dataset: Collection, Parsing, and Comparative Analysis
topic Computation and Language
url https://arxiv.org/abs/2510.14128