Self Distillation via Iterative Constructive Perturbations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dave, Maheak, Singh, Aniket Kumar, Pareek, Aryan, Jha, Harshita, Chaudhuri, Debasis, Singh, Manish Pratap
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912385064763392
author Dave, Maheak
Singh, Aniket Kumar
Pareek, Aryan
Jha, Harshita
Chaudhuri, Debasis
Singh, Manish Pratap
author_facet Dave, Maheak
Singh, Aniket Kumar
Pareek, Aryan
Jha, Harshita
Chaudhuri, Debasis
Singh, Manish Pratap
contents Deep Neural Networks have achieved remarkable achievements across various domains, however balancing performance and generalization still remains a challenge while training these networks. In this paper, we propose a novel framework that uses a cyclic optimization strategy to concurrently optimize the model and its input data for better training, rethinking the traditional training paradigm. Central to our approach is Iterative Constructive Perturbation (ICP), which leverages the model's loss to iteratively perturb the input, progressively constructing an enhanced representation over some refinement steps. This ICP input is then fed back into the model to produce improved intermediate features, which serve as a target in a self-distillation framework against the original features. By alternately altering the model's parameters to the data and the data to the model, our method effectively addresses the gap between fitting and generalization, leading to enhanced performance. Extensive experiments demonstrate that our approach not only mitigates common performance bottlenecks in neural networks but also demonstrates significant improvements across training variations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self Distillation via Iterative Constructive Perturbations
Dave, Maheak
Singh, Aniket Kumar
Pareek, Aryan
Jha, Harshita
Chaudhuri, Debasis
Singh, Manish Pratap
Machine Learning
Artificial Intelligence
Emerging Technologies
Deep Neural Networks have achieved remarkable achievements across various domains, however balancing performance and generalization still remains a challenge while training these networks. In this paper, we propose a novel framework that uses a cyclic optimization strategy to concurrently optimize the model and its input data for better training, rethinking the traditional training paradigm. Central to our approach is Iterative Constructive Perturbation (ICP), which leverages the model's loss to iteratively perturb the input, progressively constructing an enhanced representation over some refinement steps. This ICP input is then fed back into the model to produce improved intermediate features, which serve as a target in a self-distillation framework against the original features. By alternately altering the model's parameters to the data and the data to the model, our method effectively addresses the gap between fitting and generalization, leading to enhanced performance. Extensive experiments demonstrate that our approach not only mitigates common performance bottlenecks in neural networks but also demonstrates significant improvements across training variations.
title Self Distillation via Iterative Constructive Perturbations
topic Machine Learning
Artificial Intelligence
Emerging Technologies
url https://arxiv.org/abs/2505.14751