Multi-resolution subsampling for large-scale linear classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Haolin, Dette, Holger, Yu, Jun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911949283917824
author Chen, Haolin
Dette, Holger
Yu, Jun
author_facet Chen, Haolin
Dette, Holger
Yu, Jun
contents Subsampling is one of the popular methods to balance statistical efficiency and computational efficiency in the big data era. Most approaches aim at selecting informative or representative sample points to achieve good overall information of the full data. The present work takes the view that sampling techniques are recommended for the region we focus on and summary measures are enough to collect the information for the rest according to a well-designed data partitioning. We propose a multi-resolution subsampling strategy that combines global information described by summary measures and local information obtained from selected subsample points. We show that the proposed method will lead to a more efficient subsample-based estimator for general large-scale classification problems. Some asymptotic properties of the proposed method are established and connections to existing subsampling procedures are explored. Finally, we illustrate the proposed subsampling strategy via simulated and real-world examples.
format Preprint
id arxiv_https___arxiv_org_abs_2407_05691
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-resolution subsampling for large-scale linear classification
Chen, Haolin
Dette, Holger
Yu, Jun
Methodology
Statistics Theory
Subsampling is one of the popular methods to balance statistical efficiency and computational efficiency in the big data era. Most approaches aim at selecting informative or representative sample points to achieve good overall information of the full data. The present work takes the view that sampling techniques are recommended for the region we focus on and summary measures are enough to collect the information for the rest according to a well-designed data partitioning. We propose a multi-resolution subsampling strategy that combines global information described by summary measures and local information obtained from selected subsample points. We show that the proposed method will lead to a more efficient subsample-based estimator for general large-scale classification problems. Some asymptotic properties of the proposed method are established and connections to existing subsampling procedures are explored. Finally, we illustrate the proposed subsampling strategy via simulated and real-world examples.
title Multi-resolution subsampling for large-scale linear classification
topic Methodology
Statistics Theory
url https://arxiv.org/abs/2407.05691