Data Augmentation via Causal-Residual Bootstrapping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gajewski, Mateusz, Xiao, Sophia, Mazaheri, Bijan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908889506643968
author Gajewski, Mateusz
Xiao, Sophia
Mazaheri, Bijan
author_facet Gajewski, Mateusz
Xiao, Sophia
Mazaheri, Bijan
contents Data augmentation integrates domain knowledge into a dataset by making domain-informed modifications to existing data points. For example, image data can be augmented by duplicating images in different tints or orientations, thereby incorporating the knowledge that images may vary in these dimensions. Recent work by Teshima and Sugiyama has explored the integration of causal knowledge (e.g, A causes B causes C) up to conditional independence equivalence. We suggest a related approach for settings with additive noise that can incorporate information beyond a Markov equivalence class. The approach, built on the principle of independent mechanisms, permutes the residuals of models built on marginal probability distributions. Predictive models built on our augmented data demonstrate improved accuracy, for which we provide theoretical backing in linear Gaussian settings.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15335
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Data Augmentation via Causal-Residual Bootstrapping
Gajewski, Mateusz
Xiao, Sophia
Mazaheri, Bijan
Machine Learning
Data augmentation integrates domain knowledge into a dataset by making domain-informed modifications to existing data points. For example, image data can be augmented by duplicating images in different tints or orientations, thereby incorporating the knowledge that images may vary in these dimensions. Recent work by Teshima and Sugiyama has explored the integration of causal knowledge (e.g, A causes B causes C) up to conditional independence equivalence. We suggest a related approach for settings with additive noise that can incorporate information beyond a Markov equivalence class. The approach, built on the principle of independent mechanisms, permutes the residuals of models built on marginal probability distributions. Predictive models built on our augmented data demonstrate improved accuracy, for which we provide theoretical backing in linear Gaussian settings.
title Data Augmentation via Causal-Residual Bootstrapping
topic Machine Learning
url https://arxiv.org/abs/2603.15335