Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Eliopoulos, Nick John, Jajal, Purvish, Davis, James C., Liu, Gaowen, Thiravathukal, George K., Lu, Yung-Hsiang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910691492888576
author Eliopoulos, Nick John
Jajal, Purvish
Davis, James C.
Liu, Gaowen
Thiravathukal, George K.
Lu, Yung-Hsiang
author_facet Eliopoulos, Nick John
Jajal, Purvish
Davis, James C.
Liu, Gaowen
Thiravathukal, George K.
Lu, Yung-Hsiang
contents This paper investigates how to efficiently deploy vision transformers on edge devices for small workloads. Recent methods reduce the latency of transformer neural networks by removing or merging tokens, with small accuracy degradation. However, these methods are not designed with edge device deployment in mind: they do not leverage information about the latency-workload trends to improve efficiency. We address this shortcoming in our work. First, we identify factors that affect ViT latency-workload relationships. Second, we determine token pruning schedule by leveraging non-linear latency-workload relationships. Third, we demonstrate a training-free, token pruning method utilizing this schedule. We show other methods may increase latency by 2-30%, while we reduce latency by 9-26%. For similar latency (within 5.2% or 7ms) across devices we achieve 78.6%-84.5% ImageNet1K accuracy, while the state-of-the-art, Token Merging, achieves 45.8%-85.4%.
format Preprint
id arxiv_https___arxiv_org_abs_2407_05941
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge
Eliopoulos, Nick John
Jajal, Purvish
Davis, James C.
Liu, Gaowen
Thiravathukal, George K.
Lu, Yung-Hsiang
Machine Learning
Computer Vision and Pattern Recognition
This paper investigates how to efficiently deploy vision transformers on edge devices for small workloads. Recent methods reduce the latency of transformer neural networks by removing or merging tokens, with small accuracy degradation. However, these methods are not designed with edge device deployment in mind: they do not leverage information about the latency-workload trends to improve efficiency. We address this shortcoming in our work. First, we identify factors that affect ViT latency-workload relationships. Second, we determine token pruning schedule by leveraging non-linear latency-workload relationships. Third, we demonstrate a training-free, token pruning method utilizing this schedule. We show other methods may increase latency by 2-30%, while we reduce latency by 9-26%. For similar latency (within 5.2% or 7ms) across devices we achieve 78.6%-84.5% ImageNet1K accuracy, while the state-of-the-art, Token Merging, achieves 45.8%-85.4%.
title Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.05941