AutoGMM: Automatic Gaussian Mixture Modeling in Python
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2019
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909776901832704 |
|---|---|
| author | Liu, Tingshan Athey, Thomas L. Pedigo, Benjamin D. Vogelstein, Joshua T. |
| author_facet | Liu, Tingshan Athey, Thomas L. Pedigo, Benjamin D. Vogelstein, Joshua T. |
| contents | The exponential growth of complex data demands fully automatic clustering. Gaussian mixture models (GMMs) provide uncertainty-aware grouping but often require expertise to specify hyperparameters, e.g., component count and covariance structure. While mclust (R) automates this via Bayesian Information Criterion (BIC), Python lacks a comparable tool. We introduce AutoGMM, an open-source Python package automating GMM via strategic initialization using an agglomerative Mahalanobis heuristic, and parallelized model selection by information criteria. AutoGMM is a drop-in tool that yields strong out-of-the-box performance on classic benchmarks, targeted stress tests, and two real datasets, with favorable runtime scaling. The code is available at https://github.com/neurodata/AutoGMM with tests and reproducible workflows. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_1909_02688 |
| institution | arXiv |
| publishDate | 2019 |
| record_format | arxiv |
| spellingShingle | AutoGMM: Automatic Gaussian Mixture Modeling in Python Liu, Tingshan Athey, Thomas L. Pedigo, Benjamin D. Vogelstein, Joshua T. Machine Learning The exponential growth of complex data demands fully automatic clustering. Gaussian mixture models (GMMs) provide uncertainty-aware grouping but often require expertise to specify hyperparameters, e.g., component count and covariance structure. While mclust (R) automates this via Bayesian Information Criterion (BIC), Python lacks a comparable tool. We introduce AutoGMM, an open-source Python package automating GMM via strategic initialization using an agglomerative Mahalanobis heuristic, and parallelized model selection by information criteria. AutoGMM is a drop-in tool that yields strong out-of-the-box performance on classic benchmarks, targeted stress tests, and two real datasets, with favorable runtime scaling. The code is available at https://github.com/neurodata/AutoGMM with tests and reproducible workflows. |
| title | AutoGMM: Automatic Gaussian Mixture Modeling in Python |
| topic | Machine Learning |
| url | https://arxiv.org/abs/1909.02688 |