Analysing symbolic data by pseudo-marginal methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yu, Quiroz, Matias, Beranger, Boris, Kohn, Robert, Sisson, Scott A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910090500505600
author Yang, Yu
Quiroz, Matias
Beranger, Boris
Kohn, Robert
Sisson, Scott A.
author_facet Yang, Yu
Quiroz, Matias
Beranger, Boris
Kohn, Robert
Sisson, Scott A.
contents Symbolic data analysis (SDA) aggregates large individual-level datasets into a small number of distributional summaries, such as random rectangles or random histograms. The inference is carried out using these summaries in place of the original dataset, resulting in computational gains at the loss of some information. In likelihood-based SDA, the likelihood function is characterised by an integral with a large exponent, which limits the method's utility as for typical models the integral is unavailable in closed form. In addition, the likelihood function is known to produce biased parameter estimates in some circumstances. Our article develops a Bayesian framework for SDA methods in these settings that resolves the issues resulting from integral intractability and biased parameter estimation using pseudo-marginal Markov chain Monte Carlo methods. We develop an exact but computationally expensive method based on path sampling and the Poisson estimator, and a much faster, but approximate, method based on a Taylor expansion. Through simulation and real-data examples we demonstrate the performance of the developed methods, showing large reductions in computation time compared to the full-data analysis, with only a small loss of information.
format Preprint
id arxiv_https___arxiv_org_abs_2408_04419
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Analysing symbolic data by pseudo-marginal methods
Yang, Yu
Quiroz, Matias
Beranger, Boris
Kohn, Robert
Sisson, Scott A.
Methodology
Computation
Symbolic data analysis (SDA) aggregates large individual-level datasets into a small number of distributional summaries, such as random rectangles or random histograms. The inference is carried out using these summaries in place of the original dataset, resulting in computational gains at the loss of some information. In likelihood-based SDA, the likelihood function is characterised by an integral with a large exponent, which limits the method's utility as for typical models the integral is unavailable in closed form. In addition, the likelihood function is known to produce biased parameter estimates in some circumstances. Our article develops a Bayesian framework for SDA methods in these settings that resolves the issues resulting from integral intractability and biased parameter estimation using pseudo-marginal Markov chain Monte Carlo methods. We develop an exact but computationally expensive method based on path sampling and the Poisson estimator, and a much faster, but approximate, method based on a Taylor expansion. Through simulation and real-data examples we demonstrate the performance of the developed methods, showing large reductions in computation time compared to the full-data analysis, with only a small loss of information.
title Analysing symbolic data by pseudo-marginal methods
topic Methodology
Computation
url https://arxiv.org/abs/2408.04419