Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chamakh, Linda, Szabo, Zoltan
Format: Preprint
Veröffentlicht: 2021
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2110.09516
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915033659736064
author Chamakh, Linda
Szabo, Zoltan
author_facet Chamakh, Linda
Szabo, Zoltan
contents Kernel techniques are among the most popular and flexible approaches in data science allowing to represent probability measures without loss of information under mild conditions. The resulting mapping called mean embedding gives rise to a divergence measure referred to as maximum mean discrepancy (MMD) with existing quadratic-time estimators (w.r.t. the sample size) and known convergence properties for bounded kernels. In this paper we focus on the problem of MMD estimation when the mean embedding of one of the underlying distributions is available analytically. Particularly, we consider distributions on the real line (motivated by financial applications) and prove tighter concentration for the proposed estimator under this semi-explicit setting; we also extend the result to the case of unbounded (exponential) kernel with minimax-optimal lower bounds. We demonstrate the efficiency of our approach beyond synthetic example in three real-world examples relying on one-dimensional random variables: index replication and calibration on loss-given-default ratios and on S&P 500 data.
format Preprint
id arxiv_https___arxiv_org_abs_2110_09516
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Keep it Tighter -- A Story on Analytical Mean Embeddings
Chamakh, Linda
Szabo, Zoltan
Machine Learning
Portfolio Management
62C20, 62F10, 62P05, 46E22, 62B10
G.3; I.2.6
Kernel techniques are among the most popular and flexible approaches in data science allowing to represent probability measures without loss of information under mild conditions. The resulting mapping called mean embedding gives rise to a divergence measure referred to as maximum mean discrepancy (MMD) with existing quadratic-time estimators (w.r.t. the sample size) and known convergence properties for bounded kernels. In this paper we focus on the problem of MMD estimation when the mean embedding of one of the underlying distributions is available analytically. Particularly, we consider distributions on the real line (motivated by financial applications) and prove tighter concentration for the proposed estimator under this semi-explicit setting; we also extend the result to the case of unbounded (exponential) kernel with minimax-optimal lower bounds. We demonstrate the efficiency of our approach beyond synthetic example in three real-world examples relying on one-dimensional random variables: index replication and calibration on loss-given-default ratios and on S&P 500 data.
title Keep it Tighter -- A Story on Analytical Mean Embeddings
topic Machine Learning
Portfolio Management
62C20, 62F10, 62P05, 46E22, 62B10
G.3; I.2.6
url https://arxiv.org/abs/2110.09516