Web Page Classification using LLMs for Crawling Support

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sasazawa, Yuichi, Sogawa, Yasuhiro
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912370117312512
author Sasazawa, Yuichi
Sogawa, Yasuhiro
author_facet Sasazawa, Yuichi
Sogawa, Yasuhiro
contents A web crawler is a system designed to collect web pages, and efficient crawling of new pages requires appropriate algorithms. While website features such as XML sitemaps and the frequency of past page updates provide important clues for accessing new pages, their universal application across diverse conditions is challenging. In this study, we propose a method to efficiently collect new pages by classifying web pages into two types, "Index Pages" and "Content Pages," using a large language model (LLM), and leveraging the classification results to select index pages as starting points for accessing new pages. We construct a dataset with automatically annotated web page types and evaluate our approach from two perspectives: the page type classification performance and coverage of new pages. Experimental results demonstrate that the LLM-based method outperformed baseline methods in both evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06972
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Web Page Classification using LLMs for Crawling Support
Sasazawa, Yuichi
Sogawa, Yasuhiro
Information Retrieval
Computation and Language
A web crawler is a system designed to collect web pages, and efficient crawling of new pages requires appropriate algorithms. While website features such as XML sitemaps and the frequency of past page updates provide important clues for accessing new pages, their universal application across diverse conditions is challenging. In this study, we propose a method to efficiently collect new pages by classifying web pages into two types, "Index Pages" and "Content Pages," using a large language model (LLM), and leveraging the classification results to select index pages as starting points for accessing new pages. We construct a dataset with automatically annotated web page types and evaluate our approach from two perspectives: the page type classification performance and coverage of new pages. Experimental results demonstrate that the LLM-based method outperformed baseline methods in both evaluation metrics.
title Web Page Classification using LLMs for Crawling Support
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2505.06972