Saved in:
Bibliographic Details
Main Author: Jiang, Weijie.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.10440
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909254523289600
author Jiang, Weijie.
author_facet Jiang, Weijie.
contents This paper proposes a novel model for web crawling suitable for large-scale web data acquisition. This model first divides web data into several sub-data, with each sub-data corresponding to a thread task. In each thread task, web crawling tasks are concurrently executed, and the crawled data are stored in a buffer queue, awaiting further parsing. The parsing process is also divided into several threads. By establishing the model and continuously conducting crawler tests, it is found that this model is significantly optimized compared to single-threaded approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10440
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A novel multi-threaded web crawling model
Jiang, Weijie.
Databases
This paper proposes a novel model for web crawling suitable for large-scale web data acquisition. This model first divides web data into several sub-data, with each sub-data corresponding to a thread task. In each thread task, web crawling tasks are concurrently executed, and the crawled data are stored in a buffer queue, awaiting further parsing. The parsing process is also divided into several threads. By establishing the model and continuously conducting crawler tests, it is found that this model is significantly optimized compared to single-threaded approaches.
title A novel multi-threaded web crawling model
topic Databases
url https://arxiv.org/abs/2407.10440