Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.10440 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909254523289600 |
|---|---|
| author | Jiang, Weijie. |
| author_facet | Jiang, Weijie. |
| contents | This paper proposes a novel model for web crawling suitable for large-scale web data acquisition. This model first divides web data into several sub-data, with each sub-data corresponding to a thread task. In each thread task, web crawling tasks are concurrently executed, and the crawled data are stored in a buffer queue, awaiting further parsing. The parsing process is also divided into several threads. By establishing the model and continuously conducting crawler tests, it is found that this model is significantly optimized compared to single-threaded approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_10440 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A novel multi-threaded web crawling model Jiang, Weijie. Databases This paper proposes a novel model for web crawling suitable for large-scale web data acquisition. This model first divides web data into several sub-data, with each sub-data corresponding to a thread task. In each thread task, web crawling tasks are concurrently executed, and the crawled data are stored in a buffer queue, awaiting further parsing. The parsing process is also divided into several threads. By establishing the model and continuously conducting crawler tests, it is found that this model is significantly optimized compared to single-threaded approaches. |
| title | A novel multi-threaded web crawling model |
| topic | Databases |
| url | https://arxiv.org/abs/2407.10440 |