Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bourdais, Théo, Batlle, Pau, Yang, Xianjin, Baptista, Ricardo, Rouquette, Nicolas, Owhadi, Houman
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918159499395072
author Bourdais, Théo
Batlle, Pau
Yang, Xianjin
Baptista, Ricardo
Rouquette, Nicolas
Owhadi, Houman
author_facet Bourdais, Théo
Batlle, Pau
Yang, Xianjin
Baptista, Ricardo
Rouquette, Nicolas
Owhadi, Houman
contents Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17007
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots
Bourdais, Théo
Batlle, Pau
Yang, Xianjin
Baptista, Ricardo
Rouquette, Nicolas
Owhadi, Houman
Machine Learning
Artificial Intelligence
Numerical Analysis
Social and Information Networks
62A09, 62H22, 65S05, 90C35, 94C15, 46E22, 62J02, 15A83, 62D20, 68R10
Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.
title Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots
topic Machine Learning
Artificial Intelligence
Numerical Analysis
Social and Information Networks
62A09, 62H22, 65S05, 90C35, 94C15, 46E22, 62J02, 15A83, 62D20, 68R10
url https://arxiv.org/abs/2311.17007