Saved in:
Bibliographic Details
Main Authors: Reichinger, Julian, Krismayer, Thomas, Rellermeyer, Jan
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.18355
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913285848170496
author Reichinger, Julian
Krismayer, Thomas
Rellermeyer, Jan
author_facet Reichinger, Julian
Krismayer, Thomas
Rellermeyer, Jan
contents Modern, large scale monitoring systems have to process and store vast amounts of log data in near real-time. At query time the systems have to find relevant logs based on the content of the log message using support structures that can scale to these amounts of data while still being efficient to use. We present our novel Compressed Probabilistic Retrieval algorithm (COPR), capable of answering Multi-Set Multi-Membership-Queries, that can be used as an alternative to existing indexing structures for streamed log data. In our experiments, COPR required up to 93% less storage space than the tested state-of-the-art inverted index and had up to four orders of magnitude less false-positives than the tested state-of-the-art membership sketch. Additionally, COPR achieved up to 250 times higher query throughput than the tested inverted index and up to 240 times higher query throughput than the tested membership sketch.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18355
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle COPR -- Efficient, large-scale log storage and retrieval
Reichinger, Julian
Krismayer, Thomas
Rellermeyer, Jan
Information Retrieval
Databases
Data Structures and Algorithms
H.3.1
Modern, large scale monitoring systems have to process and store vast amounts of log data in near real-time. At query time the systems have to find relevant logs based on the content of the log message using support structures that can scale to these amounts of data while still being efficient to use. We present our novel Compressed Probabilistic Retrieval algorithm (COPR), capable of answering Multi-Set Multi-Membership-Queries, that can be used as an alternative to existing indexing structures for streamed log data. In our experiments, COPR required up to 93% less storage space than the tested state-of-the-art inverted index and had up to four orders of magnitude less false-positives than the tested state-of-the-art membership sketch. Additionally, COPR achieved up to 250 times higher query throughput than the tested inverted index and up to 240 times higher query throughput than the tested membership sketch.
title COPR -- Efficient, large-scale log storage and retrieval
topic Information Retrieval
Databases
Data Structures and Algorithms
H.3.1
url https://arxiv.org/abs/2402.18355