TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shen, Chaoyao, Jiang, Linfeng, Shen, Yixian, Xu, Tao, Li, Guoqing, Pathania, Anuj, Pimentel, Andy D., Zhang, Meng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918446336311296
author Shen, Chaoyao
Jiang, Linfeng
Shen, Yixian
Xu, Tao
Li, Guoqing
Pathania, Anuj
Pimentel, Andy D.
Zhang, Meng
author_facet Shen, Chaoyao
Jiang, Linfeng
Shen, Yixian
Xu, Tao
Li, Guoqing
Pathania, Anuj
Pimentel, Andy D.
Zhang, Meng
contents Deep learning (DL) compilers rely on cost models and auto-tuning to optimize tensor programs for target hardware. However, existing approaches depend on large offline datasets, incurring high collection costs and offering suboptimal transferability across platforms. In this paper, we introduce TCL, a novel efficient and transferable compiler framework for fast tensor program optimization across diverse hardware platforms to address these challenges. Specifically, TCL is built on three core enablers: (1) the RDU Sampler, a data-efficient active learning strategy that selects only 10% of tensor programs by jointly optimizing Representativeness, Diversity, and Uncertainty, substantially reducing data collection costs while maintaining near-original model accuracy; (2) a new Mamba-based cost model that efficiently captures long-range schedule dependencies while achieving a favorable trade-off between prediction accuracy and computational cost through reduced parameterization and lightweight sequence modeling; and (3) a continuous knowledge distillation framework that effectively and progressively transfers knowledge across multiple hardware platforms while avoiding the parameter explosion and data dependency issues typically caused by traditional multi-task learning. Extensive experiments validate the effectiveness of each individual enabler and the holistic TCL framework. When optimizing a range of mainstream DL models on both CPU and GPU platforms, TCL achieves, on average, 16.8x and 12.48x faster tuning time, and 1.20x and 1.13x lower inference latency, respectively, compared to Tenset-MLP.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12891
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
Shen, Chaoyao
Jiang, Linfeng
Shen, Yixian
Xu, Tao
Li, Guoqing
Pathania, Anuj
Pimentel, Andy D.
Zhang, Meng
Machine Learning
Hardware Architecture
Deep learning (DL) compilers rely on cost models and auto-tuning to optimize tensor programs for target hardware. However, existing approaches depend on large offline datasets, incurring high collection costs and offering suboptimal transferability across platforms. In this paper, we introduce TCL, a novel efficient and transferable compiler framework for fast tensor program optimization across diverse hardware platforms to address these challenges. Specifically, TCL is built on three core enablers: (1) the RDU Sampler, a data-efficient active learning strategy that selects only 10% of tensor programs by jointly optimizing Representativeness, Diversity, and Uncertainty, substantially reducing data collection costs while maintaining near-original model accuracy; (2) a new Mamba-based cost model that efficiently captures long-range schedule dependencies while achieving a favorable trade-off between prediction accuracy and computational cost through reduced parameterization and lightweight sequence modeling; and (3) a continuous knowledge distillation framework that effectively and progressively transfers knowledge across multiple hardware platforms while avoiding the parameter explosion and data dependency issues typically caused by traditional multi-task learning. Extensive experiments validate the effectiveness of each individual enabler and the holistic TCL framework. When optimizing a range of mainstream DL models on both CPU and GPU platforms, TCL achieves, on average, 16.8x and 12.48x faster tuning time, and 1.20x and 1.13x lower inference latency, respectively, compared to Tenset-MLP.
title TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2604.12891