Grounding and Enhancing Informativeness and Utility in Dataset Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shaobo, Yang, Yantai, Chen, Guo, Li, Peiru, Li, Kaixin, Zhou, Yufa, Chen, Zhaorun, Zhang, Linfeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917231105933312
author Wang, Shaobo
Yang, Yantai
Chen, Guo
Li, Peiru
Li, Kaixin
Zhou, Yufa
Chen, Zhaorun
Zhang, Linfeng
author_facet Wang, Shaobo
Yang, Yantai
Chen, Guo
Li, Peiru
Li, Kaixin
Zhou, Yufa
Chen, Zhaorun
Zhang, Linfeng
contents Dataset Distillation (DD) seeks to create a compact dataset from a large, real-world dataset. While recent methods often rely on heuristic approaches to balance efficiency and quality, the fundamental relationship between original and synthetic data remains underexplored. This paper revisits knowledge distillation-based dataset distillation within a solid theoretical framework. We introduce the concepts of Informativeness and Utility, capturing crucial information within a sample and essential samples in the training set, respectively. Building on these principles, we define optimal dataset distillation mathematically. We then present InfoUtil, a framework that balances informativeness and utility in synthesizing the distilled dataset. InfoUtil incorporates two key components: (1) game-theoretic informativeness maximization using Shapley Value attribution to extract key information from samples, and (2) principled utility maximization by selecting globally influential samples based on Gradient Norm. These components ensure that the distilled dataset is both informative and utility-optimized. Experiments demonstrate that our method achieves a 6.1\% performance improvement over the previous state-of-the-art approach on ImageNet-1K dataset using ResNet-18.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21296
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Grounding and Enhancing Informativeness and Utility in Dataset Distillation
Wang, Shaobo
Yang, Yantai
Chen, Guo
Li, Peiru
Li, Kaixin
Zhou, Yufa
Chen, Zhaorun
Zhang, Linfeng
Machine Learning
Artificial Intelligence
Dataset Distillation (DD) seeks to create a compact dataset from a large, real-world dataset. While recent methods often rely on heuristic approaches to balance efficiency and quality, the fundamental relationship between original and synthetic data remains underexplored. This paper revisits knowledge distillation-based dataset distillation within a solid theoretical framework. We introduce the concepts of Informativeness and Utility, capturing crucial information within a sample and essential samples in the training set, respectively. Building on these principles, we define optimal dataset distillation mathematically. We then present InfoUtil, a framework that balances informativeness and utility in synthesizing the distilled dataset. InfoUtil incorporates two key components: (1) game-theoretic informativeness maximization using Shapley Value attribution to extract key information from samples, and (2) principled utility maximization by selecting globally influential samples based on Gradient Norm. These components ensure that the distilled dataset is both informative and utility-optimized. Experiments demonstrate that our method achieves a 6.1\% performance improvement over the previous state-of-the-art approach on ImageNet-1K dataset using ResNet-18.
title Grounding and Enhancing Informativeness and Utility in Dataset Distillation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.21296