Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Chen, Liu, Shuming, Mangalam, Karttikeya, Qian, Guocheng, Zohra, Fatimah, Alghannam, Abdulmohsen, Malik, Jitendra, Ghanem, Bernard
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2401.04105
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916185430294528
author Zhao, Chen
Liu, Shuming
Mangalam, Karttikeya
Qian, Guocheng
Zohra, Fatimah
Alghannam, Abdulmohsen
Malik, Jitendra
Ghanem, Bernard
author_facet Zhao, Chen
Liu, Shuming
Mangalam, Karttikeya
Qian, Guocheng
Zohra, Fatimah
Alghannam, Abdulmohsen
Malik, Jitendra
Ghanem, Bernard
contents Large pretrained models are increasingly crucial in modern computer vision tasks. These models are typically used in downstream tasks by end-to-end finetuning, which is highly memory-intensive for tasks with high-resolution data, e.g., video understanding, small object detection, and point cloud analysis. In this paper, we propose Dynamic Reversible Dual-Residual Networks, or Dr$^2$Net, a novel family of network architectures that acts as a surrogate network to finetune a pretrained model with substantially reduced memory consumption. Dr$^2$Net contains two types of residual connections, one maintaining the residual structure in the pretrained models, and the other making the network reversible. Due to its reversibility, intermediate activations, which can be reconstructed from output, are cleared from memory during training. We use two coefficients on either type of residual connections respectively, and introduce a dynamic training strategy that seamlessly transitions the pretrained model to a reversible network with much higher numerical precision. We evaluate Dr$^2$Net on various pretrained models and various tasks, and show that it can reach comparable performance to conventional finetuning but with significantly less memory usage.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04105
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dr$^2$Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
Zhao, Chen
Liu, Shuming
Mangalam, Karttikeya
Qian, Guocheng
Zohra, Fatimah
Alghannam, Abdulmohsen
Malik, Jitendra
Ghanem, Bernard
Computer Vision and Pattern Recognition
Artificial Intelligence
Large pretrained models are increasingly crucial in modern computer vision tasks. These models are typically used in downstream tasks by end-to-end finetuning, which is highly memory-intensive for tasks with high-resolution data, e.g., video understanding, small object detection, and point cloud analysis. In this paper, we propose Dynamic Reversible Dual-Residual Networks, or Dr$^2$Net, a novel family of network architectures that acts as a surrogate network to finetune a pretrained model with substantially reduced memory consumption. Dr$^2$Net contains two types of residual connections, one maintaining the residual structure in the pretrained models, and the other making the network reversible. Due to its reversibility, intermediate activations, which can be reconstructed from output, are cleared from memory during training. We use two coefficients on either type of residual connections respectively, and introduce a dynamic training strategy that seamlessly transitions the pretrained model to a reversible network with much higher numerical precision. We evaluate Dr$^2$Net on various pretrained models and various tasks, and show that it can reach comparable performance to conventional finetuning but with significantly less memory usage.
title Dr$^2$Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2401.04105