Saved in:
Bibliographic Details
Main Authors: Kikuchi, Hanato, Masuya, Ryosuke, Kawamoto, Kazuhiko, Kera, Hiroshi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.07648
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918489392939008
author Kikuchi, Hanato
Masuya, Ryosuke
Kawamoto, Kazuhiko
Kera, Hiroshi
author_facet Kikuchi, Hanato
Masuya, Ryosuke
Kawamoto, Kazuhiko
Kera, Hiroshi
contents Learning parity functions, more general modular addition, is a challenging machine learning task due to its input sensitivity. A recent study substantially scaled modular addition learning in both the number of summands and the modulus. Its key idea is to increase zeros in training sequences, reducing the effective number of summands and thus controlling training difficulty; however, this induces covariate shift between training and test input distributions. This study theoretically and empirically analyzes this side effect and proposes a covariate-shift-free method for modular addition. Specifically, we introduce an auxiliary modulus $Kq$ during training, which reduces wrap-around frequency and problem difficulty while preserving the same input distribution across training and testing. Experiments show strong scalability and sample efficiency: even for large input length $N$, large modulus $q$, and small datasets -- where the sparse method fails to learn -- our method achieves equal or better match accuracy and relaxed $τ$-accuracy. For example, at $N=64$ and $q=974269$, our method trained on 100K samples achieves $97.0\%$ $τ$-accuracy at $τ=0.05$, while the sparse method achieves only $9.5\%$ with the same data size and $93.9\%$ even when extended to 1M samples.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07648
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Large-Scale Modular Addition with an Auxiliary Modulus
Kikuchi, Hanato
Masuya, Ryosuke
Kawamoto, Kazuhiko
Kera, Hiroshi
Machine Learning
Learning parity functions, more general modular addition, is a challenging machine learning task due to its input sensitivity. A recent study substantially scaled modular addition learning in both the number of summands and the modulus. Its key idea is to increase zeros in training sequences, reducing the effective number of summands and thus controlling training difficulty; however, this induces covariate shift between training and test input distributions. This study theoretically and empirically analyzes this side effect and proposes a covariate-shift-free method for modular addition. Specifically, we introduce an auxiliary modulus $Kq$ during training, which reduces wrap-around frequency and problem difficulty while preserving the same input distribution across training and testing. Experiments show strong scalability and sample efficiency: even for large input length $N$, large modulus $q$, and small datasets -- where the sparse method fails to learn -- our method achieves equal or better match accuracy and relaxed $τ$-accuracy. For example, at $N=64$ and $q=974269$, our method trained on 100K samples achieves $97.0\%$ $τ$-accuracy at $τ=0.05$, while the sparse method achieves only $9.5\%$ with the same data size and $93.9\%$ even when extended to 1M samples.
title Learning Large-Scale Modular Addition with an Auxiliary Modulus
topic Machine Learning
url https://arxiv.org/abs/2605.07648