Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Abdi, Masoud, Nahavandi, Saeid
Format: Preprint
Publié: 2016
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910674631786496
author Abdi, Masoud
Nahavandi, Saeid
author_facet Abdi, Masoud
Nahavandi, Saeid
contents In this article, we take one step toward understanding the learning behavior of deep residual networks, and supporting the observation that deep residual networks behave like ensembles. We propose a new convolutional neural network architecture which builds upon the success of residual networks by explicitly exploiting the interpretation of very deep networks as an ensemble. The proposed multi-residual network increases the number of residual functions in the residual blocks. Our architecture generates models that are wider, rather than deeper, which significantly improves accuracy. We show that our model achieves an error rate of 3.73% and 19.45% on CIFAR-10 and CIFAR-100 respectively, that outperforms almost all of the existing models. We also demonstrate that our model outperforms very deep residual networks by 0.22% (top-1 error) on the full ImageNet 2012 classification dataset. Additionally, inspired by the parallel structure of multi-residual networks, a model parallelism technique has been investigated. The model parallelism method distributes the computation of residual blocks among the processors, yielding up to 15% computational complexity improvement.
format Preprint
id arxiv_https___arxiv_org_abs_1609_05672
institution arXiv
publishDate 2016
record_format arxiv
spellingShingle Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks
Abdi, Masoud
Nahavandi, Saeid
Computer Vision and Pattern Recognition
In this article, we take one step toward understanding the learning behavior of deep residual networks, and supporting the observation that deep residual networks behave like ensembles. We propose a new convolutional neural network architecture which builds upon the success of residual networks by explicitly exploiting the interpretation of very deep networks as an ensemble. The proposed multi-residual network increases the number of residual functions in the residual blocks. Our architecture generates models that are wider, rather than deeper, which significantly improves accuracy. We show that our model achieves an error rate of 3.73% and 19.45% on CIFAR-10 and CIFAR-100 respectively, that outperforms almost all of the existing models. We also demonstrate that our model outperforms very deep residual networks by 0.22% (top-1 error) on the full ImageNet 2012 classification dataset. Additionally, inspired by the parallel structure of multi-residual networks, a model parallelism technique has been investigated. The model parallelism method distributes the computation of residual blocks among the processors, yielding up to 15% computational complexity improvement.
title Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/1609.05672