DANCE: DAta-Network Co-optimization for Efficient Segmentation Model Training and Inference

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Chaojian, Chen, Wuyang, Gu, Yuchen, Chen, Tianlong, Fu, Yonggan, Wang, Zhangyang, Lin, Yingyan Celine
Formato: Preprint
Publicado: 2021
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910896807215104
author Li, Chaojian
Chen, Wuyang
Gu, Yuchen
Chen, Tianlong
Fu, Yonggan
Wang, Zhangyang
Lin, Yingyan Celine
author_facet Li, Chaojian
Chen, Wuyang
Gu, Yuchen
Chen, Tianlong
Fu, Yonggan
Wang, Zhangyang
Lin, Yingyan Celine
contents Semantic segmentation for scene understanding is nowadays widely demanded, raising significant challenges for the algorithm efficiency, especially its applications on resource-limited platforms. Current segmentation models are trained and evaluated on massive high-resolution scene images ("data level") and suffer from the expensive computation arising from the required multi-scale aggregation("network level"). In both folds, the computational and energy costs in training and inference are notable due to the often desired large input resolutions and heavy computational burden of segmentation models. To this end, we propose DANCE, general automated DAta-Network Co-optimization for Efficient segmentation model training and inference. Distinct from existing efficient segmentation approaches that focus merely on light-weight network design, DANCE distinguishes itself as an automated simultaneous data-network co-optimization via both input data manipulation and network architecture slimming. Specifically, DANCE integrates automated data slimming which adaptively downsamples/drops input images and controls their corresponding contribution to the training loss guided by the images' spatial complexity. Such a downsampling operation, in addition to slimming down the cost associated with the input size directly, also shrinks the dynamic range of input object and context scales, therefore motivating us to also adaptively slim the network to match the downsampled data. Extensive experiments and ablating studies (on four SOTA segmentation models with three popular segmentation datasets under two training settings) demonstrate that DANCE can achieve "all-win" towards efficient segmentation(reduced training cost, less expensive inference, and better mean Intersection-over-Union (mIoU)).
format Preprint
id arxiv_https___arxiv_org_abs_2107_07706
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle DANCE: DAta-Network Co-optimization for Efficient Segmentation Model Training and Inference
Li, Chaojian
Chen, Wuyang
Gu, Yuchen
Chen, Tianlong
Fu, Yonggan
Wang, Zhangyang
Lin, Yingyan Celine
Computer Vision and Pattern Recognition
Machine Learning
Semantic segmentation for scene understanding is nowadays widely demanded, raising significant challenges for the algorithm efficiency, especially its applications on resource-limited platforms. Current segmentation models are trained and evaluated on massive high-resolution scene images ("data level") and suffer from the expensive computation arising from the required multi-scale aggregation("network level"). In both folds, the computational and energy costs in training and inference are notable due to the often desired large input resolutions and heavy computational burden of segmentation models. To this end, we propose DANCE, general automated DAta-Network Co-optimization for Efficient segmentation model training and inference. Distinct from existing efficient segmentation approaches that focus merely on light-weight network design, DANCE distinguishes itself as an automated simultaneous data-network co-optimization via both input data manipulation and network architecture slimming. Specifically, DANCE integrates automated data slimming which adaptively downsamples/drops input images and controls their corresponding contribution to the training loss guided by the images' spatial complexity. Such a downsampling operation, in addition to slimming down the cost associated with the input size directly, also shrinks the dynamic range of input object and context scales, therefore motivating us to also adaptively slim the network to match the downsampled data. Extensive experiments and ablating studies (on four SOTA segmentation models with three popular segmentation datasets under two training settings) demonstrate that DANCE can achieve "all-win" towards efficient segmentation(reduced training cost, less expensive inference, and better mean Intersection-over-Union (mIoU)).
title DANCE: DAta-Network Co-optimization for Efficient Segmentation Model Training and Inference
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2107.07706