Estimating Coverage in Streams via a Modified CVM Method
Fuente:
arXiv
Guardado en:
| Autor principal: | |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915231492472832 |
|---|---|
| author | Hernandez-Suarez, Carlos |
| author_facet | Hernandez-Suarez, Carlos |
| contents | When individuals in a population can be classified in classes or categories, the coverage of a sample, $C$, is defined as the probability that a randomly selected individual from the population belongs to a class represented in the sample. Estimating coverage is challenging because $C$ is not a fixed population parameter, but a property of the sample, and the task becomes more complex when the number of classes is unknown. Furthermore, this problem has not been addressed in scenarios where data arrive as a stream, under the constraint that only $n$ elements can be stored at a time. In this paper, we propose a simple and efficient method to estimate $C$ in streaming settings, based on a straightforward modification of the CVM algorithm, which is commonly used to estimate the number of distinct elements in a data stream. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_04567 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Estimating Coverage in Streams via a Modified CVM Method Hernandez-Suarez, Carlos Computation Machine Learning 68W20 When individuals in a population can be classified in classes or categories, the coverage of a sample, $C$, is defined as the probability that a randomly selected individual from the population belongs to a class represented in the sample. Estimating coverage is challenging because $C$ is not a fixed population parameter, but a property of the sample, and the task becomes more complex when the number of classes is unknown. Furthermore, this problem has not been addressed in scenarios where data arrive as a stream, under the constraint that only $n$ elements can be stored at a time. In this paper, we propose a simple and efficient method to estimate $C$ in streaming settings, based on a straightforward modification of the CVM algorithm, which is commonly used to estimate the number of distinct elements in a data stream. |
| title | Estimating Coverage in Streams via a Modified CVM Method |
| topic | Computation Machine Learning 68W20 |
| url | https://arxiv.org/abs/2504.04567 |