DRUPI: Dataset Reduction Using Privileged Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shaobo, Jiang, Youxin, Niu, Tianle, Yang, Yantai, Zhang, Ruiji, Hu, Shuhao, Zhang, Shuaiyu, Sun, Chenghao, Li, Weiya, He, Conghui, Hu, Xuming, Zhang, Linfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912956374056960
author Wang, Shaobo
Jiang, Youxin
Niu, Tianle
Yang, Yantai
Zhang, Ruiji
Hu, Shuhao
Zhang, Shuaiyu
Sun, Chenghao
Li, Weiya
He, Conghui
Hu, Xuming
Zhang, Linfeng
author_facet Wang, Shaobo
Jiang, Youxin
Niu, Tianle
Yang, Yantai
Zhang, Ruiji
Hu, Shuhao
Zhang, Shuaiyu
Sun, Chenghao
Li, Weiya
He, Conghui
Hu, Xuming
Zhang, Linfeng
contents Dataset Condensation (DC) seeks to select or distill samples from large datasets into smaller subsets while preserving performance on target tasks. Existing methods primarily focus on pruning or synthesizing data in the same format as the original dataset, typically being the input data and corresponding labels. However, in DC settings, we find it is possible to synthesize more information beyond the data-label pair as an additional learning target to facilitate model training. In this paper, we introduce Dataset Condensation using Privileged Information (DCPI), which enriches DC by synthesizing privileged information alongside the reduced dataset. This privileged information can take the form of feature labels or attention labels, providing auxiliary supervision to improve model learning. Our findings reveal that effective feature labels must balance between being overly discriminative and excessively diverse, with a moderate level proves optimal for improving the reduced dataset's efficacy. Extensive experiments on ImageNet-1K, CIFAR-10/100 and Tiny ImageNet demonstrate that DCPI integrates seamlessly with existing dataset condensation methods, offering significant performance gains.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DRUPI: Dataset Reduction Using Privileged Information
Wang, Shaobo
Jiang, Youxin
Niu, Tianle
Yang, Yantai
Zhang, Ruiji
Hu, Shuhao
Zhang, Shuaiyu
Sun, Chenghao
Li, Weiya
He, Conghui
Hu, Xuming
Zhang, Linfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Dataset Condensation (DC) seeks to select or distill samples from large datasets into smaller subsets while preserving performance on target tasks. Existing methods primarily focus on pruning or synthesizing data in the same format as the original dataset, typically being the input data and corresponding labels. However, in DC settings, we find it is possible to synthesize more information beyond the data-label pair as an additional learning target to facilitate model training. In this paper, we introduce Dataset Condensation using Privileged Information (DCPI), which enriches DC by synthesizing privileged information alongside the reduced dataset. This privileged information can take the form of feature labels or attention labels, providing auxiliary supervision to improve model learning. Our findings reveal that effective feature labels must balance between being overly discriminative and excessively diverse, with a moderate level proves optimal for improving the reduced dataset's efficacy. Extensive experiments on ImageNet-1K, CIFAR-10/100 and Tiny ImageNet demonstrate that DCPI integrates seamlessly with existing dataset condensation methods, offering significant performance gains.
title DRUPI: Dataset Reduction Using Privileged Information
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.01611