To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lawrence, Hannah, Hofgard, Elyssa, Portilheiro, Vasco, Chen, Yuxuan, Smidt, Tess, Walters, Robin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915899980644352
author Lawrence, Hannah
Hofgard, Elyssa
Portilheiro, Vasco
Chen, Yuxuan
Smidt, Tess
Walters, Robin
author_facet Lawrence, Hannah
Hofgard, Elyssa
Portilheiro, Vasco
Chen, Yuxuan
Smidt, Tess
Walters, Robin
contents Symmetry-aware methods for machine learning, such as data augmentation and equivariant architectures, encourage correct model behavior on all transformations (e.g. rotations or permutations) of the original dataset. These methods can improve generalization and sample efficiency, under the assumption that the transformed datapoints are highly probable, or "important", under the test distribution. In this work, we develop a method for critically evaluating this assumption. In particular, we propose a metric to quantify the amount of symmetry breaking in a dataset, via a two-sample classifier test that distinguishes between the original dataset and its randomly augmented equivalent. We validate our metric on synthetic datasets, and then use it to uncover surprisingly high degrees of symmetry-breaking in several benchmark point cloud datasets, constituting a severe form of dataset bias. We show theoretically that distributional symmetry-breaking can prevent invariant methods from performing optimally even when the underlying labels are truly invariant, for invariant ridge regression in the infinite feature limit. Empirically, the implication for symmetry-aware methods is dataset-dependent: equivariant methods still impart benefits on some symmetry-biased datasets, but not others, particularly when the symmetry bias is predictive of the labels. Overall, these findings suggest that understanding equivariance -- both when it works, and why -- may require rethinking symmetry biases in the data.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking
Lawrence, Hannah
Hofgard, Elyssa
Portilheiro, Vasco
Chen, Yuxuan
Smidt, Tess
Walters, Robin
Machine Learning
Symmetry-aware methods for machine learning, such as data augmentation and equivariant architectures, encourage correct model behavior on all transformations (e.g. rotations or permutations) of the original dataset. These methods can improve generalization and sample efficiency, under the assumption that the transformed datapoints are highly probable, or "important", under the test distribution. In this work, we develop a method for critically evaluating this assumption. In particular, we propose a metric to quantify the amount of symmetry breaking in a dataset, via a two-sample classifier test that distinguishes between the original dataset and its randomly augmented equivalent. We validate our metric on synthetic datasets, and then use it to uncover surprisingly high degrees of symmetry-breaking in several benchmark point cloud datasets, constituting a severe form of dataset bias. We show theoretically that distributional symmetry-breaking can prevent invariant methods from performing optimally even when the underlying labels are truly invariant, for invariant ridge regression in the infinite feature limit. Empirically, the implication for symmetry-aware methods is dataset-dependent: equivariant methods still impart benefits on some symmetry-biased datasets, but not others, particularly when the symmetry bias is predictive of the labels. Overall, these findings suggest that understanding equivariance -- both when it works, and why -- may require rethinking symmetry biases in the data.
title To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking
topic Machine Learning
url https://arxiv.org/abs/2510.01349