Learning from Uncertain Data: From Possible Worlds to Possible Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jiongli, Feng, Su, Glavic, Boris, Salimi, Babak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917677726957568
author Zhu, Jiongli
Feng, Su
Glavic, Boris
Salimi, Babak
author_facet Zhu, Jiongli
Feng, Su
Glavic, Boris
Salimi, Babak
contents We introduce an efficient method for learning linear models from uncertain data, where uncertainty is represented as a set of possible variations in the data, leading to predictive multiplicity. Our approach leverages abstract interpretation and zonotopes, a type of convex polytope, to compactly represent these dataset variations, enabling the symbolic execution of gradient descent on all possible worlds simultaneously. We develop techniques to ensure that this process converges to a fixed point and derive closed-form solutions for this fixed point. Our method provides sound over-approximations of all possible optimal models and viable prediction ranges. We demonstrate the effectiveness of our approach through theoretical and empirical analysis, highlighting its potential to reason about model and prediction uncertainty due to data quality issues in training data.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning from Uncertain Data: From Possible Worlds to Possible Models
Zhu, Jiongli
Feng, Su
Glavic, Boris
Salimi, Babak
Machine Learning
Databases
Symbolic Computation
We introduce an efficient method for learning linear models from uncertain data, where uncertainty is represented as a set of possible variations in the data, leading to predictive multiplicity. Our approach leverages abstract interpretation and zonotopes, a type of convex polytope, to compactly represent these dataset variations, enabling the symbolic execution of gradient descent on all possible worlds simultaneously. We develop techniques to ensure that this process converges to a fixed point and derive closed-form solutions for this fixed point. Our method provides sound over-approximations of all possible optimal models and viable prediction ranges. We demonstrate the effectiveness of our approach through theoretical and empirical analysis, highlighting its potential to reason about model and prediction uncertainty due to data quality issues in training data.
title Learning from Uncertain Data: From Possible Worlds to Possible Models
topic Machine Learning
Databases
Symbolic Computation
url https://arxiv.org/abs/2405.18549