AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Feiyang, Chang, Nadine, Shen, Maying, Law, Marc T., Mahmood, Rafid, Jia, Ruoxi, Alvarez, Jose M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909668010360832
author Kang, Feiyang
Chang, Nadine
Shen, Maying
Law, Marc T.
Mahmood, Rafid
Jia, Ruoxi
Alvarez, Jose M.
author_facet Kang, Feiyang
Chang, Nadine
Shen, Maying
Law, Marc T.
Mahmood, Rafid
Jia, Ruoxi
Alvarez, Jose M.
contents The computational burden and inherent redundancy of large-scale datasets challenge the training of contemporary machine learning models. Data pruning offers a solution by selecting smaller, informative subsets, yet existing methods struggle: density-based approaches can be task-agnostic, while model-based techniques may introduce redundancy or prove computationally prohibitive. We introduce Adaptive De-Duplication (AdaDeDup), a novel hybrid framework that synergistically integrates density-based pruning with model-informed feedback in a cluster-adaptive manner. AdaDeDup first partitions data and applies an initial density-based pruning. It then employs a proxy model to evaluate the impact of this initial pruning within each cluster by comparing losses on kept versus pruned samples. This task-aware signal adaptively adjusts cluster-specific pruning thresholds, enabling more aggressive pruning in redundant clusters while preserving critical data in informative ones. Extensive experiments on large-scale object detection benchmarks (Waymo, COCO, nuScenes) using standard models (BEVFormer, Faster R-CNN) demonstrate AdaDeDup's advantages. It significantly outperforms prominent baselines, substantially reduces performance degradation (e.g., over 54% versus random sampling on Waymo), and achieves near-original model performance while pruning 20% of data, highlighting its efficacy in enhancing data efficiency for large-scale model training. Code is open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
Kang, Feiyang
Chang, Nadine
Shen, Maying
Law, Marc T.
Mahmood, Rafid
Jia, Ruoxi
Alvarez, Jose M.
Computer Vision and Pattern Recognition
Machine Learning
The computational burden and inherent redundancy of large-scale datasets challenge the training of contemporary machine learning models. Data pruning offers a solution by selecting smaller, informative subsets, yet existing methods struggle: density-based approaches can be task-agnostic, while model-based techniques may introduce redundancy or prove computationally prohibitive. We introduce Adaptive De-Duplication (AdaDeDup), a novel hybrid framework that synergistically integrates density-based pruning with model-informed feedback in a cluster-adaptive manner. AdaDeDup first partitions data and applies an initial density-based pruning. It then employs a proxy model to evaluate the impact of this initial pruning within each cluster by comparing losses on kept versus pruned samples. This task-aware signal adaptively adjusts cluster-specific pruning thresholds, enabling more aggressive pruning in redundant clusters while preserving critical data in informative ones. Extensive experiments on large-scale object detection benchmarks (Waymo, COCO, nuScenes) using standard models (BEVFormer, Faster R-CNN) demonstrate AdaDeDup's advantages. It significantly outperforms prominent baselines, substantially reduces performance degradation (e.g., over 54% versus random sampling on Waymo), and achieves near-original model performance while pruning 20% of data, highlighting its efficacy in enhancing data efficiency for large-scale model training. Code is open-sourced.
title AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.00049