Dataset Growth

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Ziheng, Xu, Zhaopan, Zhou, Yukun, Zheng, Zangwei, Cheng, Zebang, Tang, Hao, Shang, Lei, Sun, Baigui, Peng, Xiaojiang, Timofte, Radu, Yao, Hongxun, Wang, Kai, You, Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929431594926080
author Qin, Ziheng
Xu, Zhaopan
Zhou, Yukun
Zheng, Zangwei
Cheng, Zebang
Tang, Hao
Shang, Lei
Sun, Baigui
Peng, Xiaojiang
Timofte, Radu
Yao, Hongxun
Wang, Kai
You, Yang
author_facet Qin, Ziheng
Xu, Zhaopan
Zhou, Yukun
Zheng, Zangwei
Cheng, Zebang
Tang, Hao
Shang, Lei
Sun, Baigui
Peng, Xiaojiang
Timofte, Radu
Yao, Hongxun
Wang, Kai
You, Yang
contents Deep learning benefits from the growing abundance of available data. Meanwhile, efficiently dealing with the growing data scale has become a challenge. Data publicly available are from different sources with various qualities, and it is impractical to do manual cleaning against noise and redundancy given today's data scale. There are existing techniques for cleaning/selecting the collected data. However, these methods are mainly proposed for offline settings that target one of the cleanness and redundancy problems. In practice, data are growing exponentially with both problems. This leads to repeated data curation with sub-optimal efficiency. To tackle this challenge, we propose InfoGrowth, an efficient online algorithm for data cleaning and selection, resulting in a growing dataset that keeps up to date with awareness of cleanliness and diversity. InfoGrowth can improve data quality/efficiency on both single-modal and multi-modal tasks, with an efficient and scalable design. Its framework makes it practical for real-world data engines.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18347
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dataset Growth
Qin, Ziheng
Xu, Zhaopan
Zhou, Yukun
Zheng, Zangwei
Cheng, Zebang
Tang, Hao
Shang, Lei
Sun, Baigui
Peng, Xiaojiang
Timofte, Radu
Yao, Hongxun
Wang, Kai
You, Yang
Machine Learning
Deep learning benefits from the growing abundance of available data. Meanwhile, efficiently dealing with the growing data scale has become a challenge. Data publicly available are from different sources with various qualities, and it is impractical to do manual cleaning against noise and redundancy given today's data scale. There are existing techniques for cleaning/selecting the collected data. However, these methods are mainly proposed for offline settings that target one of the cleanness and redundancy problems. In practice, data are growing exponentially with both problems. This leads to repeated data curation with sub-optimal efficiency. To tackle this challenge, we propose InfoGrowth, an efficient online algorithm for data cleaning and selection, resulting in a growing dataset that keeps up to date with awareness of cleanliness and diversity. InfoGrowth can improve data quality/efficiency on both single-modal and multi-modal tasks, with an efficient and scalable design. Its framework makes it practical for real-world data engines.
title Dataset Growth
topic Machine Learning
url https://arxiv.org/abs/2405.18347