Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bee, Marco, Santi, Flavio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908382923849728
author Bee, Marco
Santi, Flavio
author_facet Bee, Marco
Santi, Flavio
contents We develop an unsupervised mixture model for non-negative, skewed and heavy-tailed data, such as losses in actuarial and risk management applications. The mixture has a lognormal component, which is usually appropriate for the body of the distribution, and a Pareto-type tail, aimed at accommodating the largest observations, since the lognormal tail often decays too fast. We show that maximum likelihood estimation can be performed by means of the EM algorithm and that the model is quite flexible in fitting data from different data-generating processes. Simulation experiments and a real-data application to automobiles claims suggest that the approach is equivalent in terms of goodness-of-fit, but easier to estimate, with respect to two existing distributions with similar features.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models
Bee, Marco
Santi, Flavio
Methodology
We develop an unsupervised mixture model for non-negative, skewed and heavy-tailed data, such as losses in actuarial and risk management applications. The mixture has a lognormal component, which is usually appropriate for the body of the distribution, and a Pareto-type tail, aimed at accommodating the largest observations, since the lognormal tail often decays too fast. We show that maximum likelihood estimation can be performed by means of the EM algorithm and that the model is quite flexible in fitting data from different data-generating processes. Simulation experiments and a real-data application to automobiles claims suggest that the approach is equivalent in terms of goodness-of-fit, but easier to estimate, with respect to two existing distributions with similar features.
title Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models
topic Methodology
url https://arxiv.org/abs/2505.22507