Dataset Distillation via Vision-Language Category Prototype

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zou, Yawen, Li, Guang, Su, Duo, Wang, Zi, Yu, Jun, Zhang, Chao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913919130402816
author Zou, Yawen
Li, Guang
Su, Duo
Wang, Zi
Yu, Jun
Zhang, Chao
author_facet Zou, Yawen
Li, Guang
Su, Duo
Wang, Zi
Yu, Jun
Zhang, Chao
contents Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consumption. However, previous DD methods mainly focus on distilling information from images, often overlooking the semantic information inherent in the data. The disregard for context hinders the model's generalization ability, particularly in tasks involving complex datasets, which may result in illogical outputs or the omission of critical objects. In this study, we integrate vision-language methods into DD by introducing text prototypes to distill language information and collaboratively synthesize data with image prototypes, thereby enhancing dataset distillation performance. Notably, the text prototypes utilized in this study are derived from descriptive text information generated by an open-source large language model. This framework demonstrates broad applicability across datasets without pre-existing text descriptions, expanding the potential of dataset distillation beyond traditional image-based approaches. Compared to other methods, the proposed approach generates logically coherent images containing target objects, achieving state-of-the-art validation performance and demonstrating robust generalization. Source code and generated data are available in https://github.com/zou-yawen/Dataset-Distillation-via-Vision-Language-Category-Prototype/
format Preprint
id arxiv_https___arxiv_org_abs_2506_23580
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dataset Distillation via Vision-Language Category Prototype
Zou, Yawen
Li, Guang
Su, Duo
Wang, Zi
Yu, Jun
Zhang, Chao
Computer Vision and Pattern Recognition
Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consumption. However, previous DD methods mainly focus on distilling information from images, often overlooking the semantic information inherent in the data. The disregard for context hinders the model's generalization ability, particularly in tasks involving complex datasets, which may result in illogical outputs or the omission of critical objects. In this study, we integrate vision-language methods into DD by introducing text prototypes to distill language information and collaboratively synthesize data with image prototypes, thereby enhancing dataset distillation performance. Notably, the text prototypes utilized in this study are derived from descriptive text information generated by an open-source large language model. This framework demonstrates broad applicability across datasets without pre-existing text descriptions, expanding the potential of dataset distillation beyond traditional image-based approaches. Compared to other methods, the proposed approach generates logically coherent images containing target objects, achieving state-of-the-art validation performance and demonstrating robust generalization. Source code and generated data are available in https://github.com/zou-yawen/Dataset-Distillation-via-Vision-Language-Category-Prototype/
title Dataset Distillation via Vision-Language Category Prototype
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23580