Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Yangjie, Zhu, Honglin, Qiu, Qian, Cui, Weihao, Liu, Zihan, Guo, Cong, Feng, Siyuan, Meng, Jintao, Lan, Haidong, Leng, Jingwen, Zhu, Wenxi, Deng, Minwen
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912010850009088
author Zhou, Yangjie
Zhu, Honglin
Qiu, Qian
Cui, Weihao
Liu, Zihan
Guo, Cong
Feng, Siyuan
Meng, Jintao
Lan, Haidong
Leng, Jingwen
Zhu, Wenxi
Deng, Minwen
author_facet Zhou, Yangjie
Zhu, Honglin
Qiu, Qian
Cui, Weihao
Liu, Zihan
Guo, Cong
Feng, Siyuan
Meng, Jintao
Lan, Haidong
Leng, Jingwen
Zhu, Wenxi
Deng, Minwen
contents Dynamic-shape deep neural networks (DNNs) are rapidly evolving, attracting attention for their ability to handle variable input sizes in real-time applications. However, existing compilation optimization methods for such networks often rely heavily on predefined samples to guide the compilation process, which restricts their adaptability and efficiency. These sample-driven methods struggle to efficiently manage the diverse and unpredictable shapes encountered in real-world scenarios, often resulting in suboptimal performance. To tackle these issues, we introduce Vortex, a hardware-driven and sample-free compiler tailored for dynamic-shape tensor programs. Vortex capitalizes on detailed hardware information and hierarchizes the strategy space to facilitate high-performance code generation without relying on runtime shape samples. It features a unique bidirectional compilation workflow, combining top-down abstraction for aligning tensor program execution with hardware hierarchies and bottom-up kernel construction to narrow the search space, enabling Vortex to achieve remarkable efficiency. Comprehensive evaluations confirm that Vortex reduces compilation time by $176\times$ compared to the existing dynamic-shape compiler. Additionally, it substantially outperforms existing vendor-provided libraries and dynamic-shape compilers on both CPU and GPU platforms, delivering speedups of $2.53\times$ and $3.01\times$, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01075
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
Zhou, Yangjie
Zhu, Honglin
Qiu, Qian
Cui, Weihao
Liu, Zihan
Guo, Cong
Feng, Siyuan
Meng, Jintao
Lan, Haidong
Leng, Jingwen
Zhu, Wenxi
Deng, Minwen
Distributed, Parallel, and Cluster Computing
Dynamic-shape deep neural networks (DNNs) are rapidly evolving, attracting attention for their ability to handle variable input sizes in real-time applications. However, existing compilation optimization methods for such networks often rely heavily on predefined samples to guide the compilation process, which restricts their adaptability and efficiency. These sample-driven methods struggle to efficiently manage the diverse and unpredictable shapes encountered in real-world scenarios, often resulting in suboptimal performance. To tackle these issues, we introduce Vortex, a hardware-driven and sample-free compiler tailored for dynamic-shape tensor programs. Vortex capitalizes on detailed hardware information and hierarchizes the strategy space to facilitate high-performance code generation without relying on runtime shape samples. It features a unique bidirectional compilation workflow, combining top-down abstraction for aligning tensor program execution with hardware hierarchies and bottom-up kernel construction to narrow the search space, enabling Vortex to achieve remarkable efficiency. Comprehensive evaluations confirm that Vortex reduces compilation time by $176\times$ compared to the existing dynamic-shape compiler. Additionally, it substantially outperforms existing vendor-provided libraries and dynamic-shape compilers on both CPU and GPU platforms, delivering speedups of $2.53\times$ and $3.01\times$, respectively.
title Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2409.01075