A fast and effective kernel two-sample test for large-scale data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Song, Hoseung, Chen, Hao
Natura: Preprint
Pubblicazione: 2021
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918152589279232
author Song, Hoseung
Chen, Hao
author_facet Song, Hoseung
Chen, Hao
contents Kernel two-sample tests have been widely used, and the development of efficient methods for high-dimensional, large-scale data is receiving increasing attention in the big data era. However, existing methods, such as the maximum mean discrepancy (MMD) and recently proposed kernel-based tests for large-scale data, are computationally intensive and/or ineffective for some common alternatives in high-dimensional data. In this paper, we propose a new test that exhibits high power across a wide range of alternatives. Furthermore, the new test is more robust to high dimensions than existing methods and does not require optimization procedures for choosing kernel bandwidth and other parameters through data splitting. Numerical studies demonstrate that the new approach performs well on both synthetic and real-world data.
format Preprint
id arxiv_https___arxiv_org_abs_2110_03118
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle A fast and effective kernel two-sample test for large-scale data
Song, Hoseung
Chen, Hao
Methodology
Machine Learning
Kernel two-sample tests have been widely used, and the development of efficient methods for high-dimensional, large-scale data is receiving increasing attention in the big data era. However, existing methods, such as the maximum mean discrepancy (MMD) and recently proposed kernel-based tests for large-scale data, are computationally intensive and/or ineffective for some common alternatives in high-dimensional data. In this paper, we propose a new test that exhibits high power across a wide range of alternatives. Furthermore, the new test is more robust to high dimensions than existing methods and does not require optimization procedures for choosing kernel bandwidth and other parameters through data splitting. Numerical studies demonstrate that the new approach performs well on both synthetic and real-world data.
title A fast and effective kernel two-sample test for large-scale data
topic Methodology
Machine Learning
url https://arxiv.org/abs/2110.03118