Saved in:
Bibliographic Details
Main Authors: Li, Muquan, Gou, Hang, Zhang, Dongyang, Liang, Shuang, Xie, Xiurui, Ouyang, Deqiang, Qin, Ke
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.04838
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910009347014656
author Li, Muquan
Gou, Hang
Zhang, Dongyang
Liang, Shuang
Xie, Xiurui
Ouyang, Deqiang
Qin, Ke
author_facet Li, Muquan
Gou, Hang
Zhang, Dongyang
Liang, Shuang
Xie, Xiurui
Ouyang, Deqiang
Qin, Ke
contents The growing demand for efficient deep learning has positioned dataset distillation as a pivotal technique for compressing training dataset while preserving model performance. However, existing inner-loop optimization methods for dataset distillation typically rely on random truncation strategies, which lack flexibility and often yield suboptimal results. In this work, we observe that neural networks exhibit distinct learning dynamics across different training stages-early, middle, and late-making random truncation ineffective. To address this limitation, we propose Automatic Truncated Backpropagation Through Time (AT-BPTT), a novel framework that dynamically adapts both truncation positions and window sizes according to intrinsic gradient behavior. AT-BPTT introduces three key components: (1) a probabilistic mechanism for stage-aware timestep selection, (2) an adaptive window sizing strategy based on gradient variation, and (3) a low-rank Hessian approximation to reduce computational overhead. Extensive experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-1K show that AT-BPTT achieves state-of-the-art performance, improving accuracy by an average of 6.16% over baseline methods. Moreover, our approach accelerates inner-loop optimization by 3.9x while saving 63% memory cost.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04838
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
Li, Muquan
Gou, Hang
Zhang, Dongyang
Liang, Shuang
Xie, Xiurui
Ouyang, Deqiang
Qin, Ke
Computer Vision and Pattern Recognition
Machine Learning
The growing demand for efficient deep learning has positioned dataset distillation as a pivotal technique for compressing training dataset while preserving model performance. However, existing inner-loop optimization methods for dataset distillation typically rely on random truncation strategies, which lack flexibility and often yield suboptimal results. In this work, we observe that neural networks exhibit distinct learning dynamics across different training stages-early, middle, and late-making random truncation ineffective. To address this limitation, we propose Automatic Truncated Backpropagation Through Time (AT-BPTT), a novel framework that dynamically adapts both truncation positions and window sizes according to intrinsic gradient behavior. AT-BPTT introduces three key components: (1) a probabilistic mechanism for stage-aware timestep selection, (2) an adaptive window sizing strategy based on gradient variation, and (3) a low-rank Hessian approximation to reduce computational overhead. Extensive experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-1K show that AT-BPTT achieves state-of-the-art performance, improving accuracy by an average of 6.16% over baseline methods. Moreover, our approach accelerates inner-loop optimization by 3.9x while saving 63% memory cost.
title Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.04838