Canonicalization Leakage: How Canonical Representatives Confound Supervised Learning under Group Symmetry
Fuente:
Zenodo
Saved in:
| Main Author: | |
|---|---|
| Format: | Recurso digital |
| Language: | English |
| Published: |
Zenodo
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866901925725732864 |
|---|---|
| author | Li, Alex |
| author_facet | Li, Alex |
| contents | <p>Models trained on canonical representatives of equivalence classes under group symmetry can exploit representation artifacts rather than learning invariant structure. We propose CL-DIAG, a six-step diagnostic protocol that detects, localizes, and quantifies this "canonicalization leakage." Applied to circuit complexity prediction over 616,126 NPN equivalence classes of 5-input Boolean functions (|G| = 7,680), CL-DIAG reveals that a baseline MLP achieves Spearman r_s = 0.788 on canonical data but only r_s = 0.254 when NPN-averaged, with 0% prediction consistency. Signal decomposition shows canonical performance decomposes into classical invariant signal (r_s = 0.635), neural invariant signal (+0.142), and canonicalization leakage (+0.011). NPN augmentation at 7x recovers r_s = 0.777, exceeding the classical invariant ceiling by 14 percentage points. A matched-volume control confirms the gain is from symmetry-consistent augmentation, not generic regularization.</p> <p>v2: Figure 1 now uses actual model predictions (previously used placeholder visualization). No changes to text, results, or conclusions.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19112504 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Canonicalization Leakage: How Canonical Representatives Confound Supervised Learning under Group Symmetry Li, Alex canonicalization leakage NPN equivalence group symmetry data augmentation shortcut learning Boolean functions circuit complexity <p>Models trained on canonical representatives of equivalence classes under group symmetry can exploit representation artifacts rather than learning invariant structure. We propose CL-DIAG, a six-step diagnostic protocol that detects, localizes, and quantifies this "canonicalization leakage." Applied to circuit complexity prediction over 616,126 NPN equivalence classes of 5-input Boolean functions (|G| = 7,680), CL-DIAG reveals that a baseline MLP achieves Spearman r_s = 0.788 on canonical data but only r_s = 0.254 when NPN-averaged, with 0% prediction consistency. Signal decomposition shows canonical performance decomposes into classical invariant signal (r_s = 0.635), neural invariant signal (+0.142), and canonicalization leakage (+0.011). NPN augmentation at 7x recovers r_s = 0.777, exceeding the classical invariant ceiling by 14 percentage points. A matched-volume control confirms the gain is from symmetry-consistent augmentation, not generic regularization.</p> <p>v2: Figure 1 now uses actual model predictions (previously used placeholder visualization). No changes to text, results, or conclusions.</p> |
| title | Canonicalization Leakage: How Canonical Representatives Confound Supervised Learning under Group Symmetry |
| topic | canonicalization leakage NPN equivalence group symmetry data augmentation shortcut learning Boolean functions circuit complexity |
| url | https://doi.org/10.5281/zenodo.19112504 |