Perturbation-based Effect Measures for Compositional Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lundborg, Anton Rask, Pfister, Niklas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908385564164096
author Lundborg, Anton Rask
Pfister, Niklas
author_facet Lundborg, Anton Rask
Pfister, Niklas
contents Existing effect measures for compositional features are inadequate for many modern applications, for example, in microbiome research, since they display traits such as high-dimensionality and sparsity that can be poorly modelled with traditional parametric approaches. Further, assessing -- in an unbiased way -- how summary statistics of a composition (e.g., racial diversity) affect a response variable is not straightforward. We propose a framework based on hypothetical data perturbations which defines interpretable statistical functionals on the compositions themselves, which we call average perturbation effects. These effects naturally account for confounding that biases frequently used marginal dependence analyses. We show how average perturbation effects can be estimated efficiently by deriving a perturbation-dependent reparametrization and applying semiparametric estimation techniques. We analyze the proposed estimators empirically on simulated and semi-synthetic data and demonstrate advantages over existing techniques on data from New York schools and microbiome data.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18501
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Perturbation-based Effect Measures for Compositional Data
Lundborg, Anton Rask
Pfister, Niklas
Methodology
Statistics Theory
Machine Learning
Existing effect measures for compositional features are inadequate for many modern applications, for example, in microbiome research, since they display traits such as high-dimensionality and sparsity that can be poorly modelled with traditional parametric approaches. Further, assessing -- in an unbiased way -- how summary statistics of a composition (e.g., racial diversity) affect a response variable is not straightforward. We propose a framework based on hypothetical data perturbations which defines interpretable statistical functionals on the compositions themselves, which we call average perturbation effects. These effects naturally account for confounding that biases frequently used marginal dependence analyses. We show how average perturbation effects can be estimated efficiently by deriving a perturbation-dependent reparametrization and applying semiparametric estimation techniques. We analyze the proposed estimators empirically on simulated and semi-synthetic data and demonstrate advantages over existing techniques on data from New York schools and microbiome data.
title Perturbation-based Effect Measures for Compositional Data
topic Methodology
Statistics Theory
Machine Learning
url https://arxiv.org/abs/2311.18501