Liger Kernel: Efficient Triton Kernels for LLM Training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hsu, Pin-Lun, Dai, Yun, Kothapalli, Vignesh, Song, Qingquan, Tang, Shao, Zhu, Siyu, Shimizu, Steven, Sahni, Shivam, Ning, Haowen, Chen, Yanning
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910796984877056
author Hsu, Pin-Lun
Dai, Yun
Kothapalli, Vignesh
Song, Qingquan
Tang, Shao
Zhu, Siyu
Shimizu, Steven
Sahni, Shivam
Ning, Haowen
Chen, Yanning
author_facet Hsu, Pin-Lun
Dai, Yun
Kothapalli, Vignesh
Song, Qingquan
Tang, Shao
Zhu, Siyu
Shimizu, Steven
Sahni, Shivam
Ning, Haowen
Chen, Yanning
contents Training Large Language Models (LLMs) efficiently at scale presents a formidable challenge, driven by their ever-increasing computational demands and the need for enhanced performance. In this work, we introduce Liger-Kernel, an open-sourced set of Triton kernels developed specifically for LLM training. With kernel optimization techniques like kernel operation fusing and input chunking, our kernels achieve on average a 20% increase in training throughput and a 60% reduction in GPU memory usage for popular LLMs compared to HuggingFace implementations. In addition, Liger-Kernel is designed with modularity, accessibility, and adaptability in mind, catering to both casual and expert users. Comprehensive benchmarks and integration tests are built in to ensure compatibility, performance, correctness, and convergence across diverse computing environments and model architectures. The source code is available under a permissive license at: github.com/linkedin/Liger-Kernel.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10989
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Liger Kernel: Efficient Triton Kernels for LLM Training
Hsu, Pin-Lun
Dai, Yun
Kothapalli, Vignesh
Song, Qingquan
Tang, Shao
Zhu, Siyu
Shimizu, Steven
Sahni, Shivam
Ning, Haowen
Chen, Yanning
Machine Learning
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
Training Large Language Models (LLMs) efficiently at scale presents a formidable challenge, driven by their ever-increasing computational demands and the need for enhanced performance. In this work, we introduce Liger-Kernel, an open-sourced set of Triton kernels developed specifically for LLM training. With kernel optimization techniques like kernel operation fusing and input chunking, our kernels achieve on average a 20% increase in training throughput and a 60% reduction in GPU memory usage for popular LLMs compared to HuggingFace implementations. In addition, Liger-Kernel is designed with modularity, accessibility, and adaptability in mind, catering to both casual and expert users. Comprehensive benchmarks and integration tests are built in to ensure compatibility, performance, correctness, and convergence across diverse computing environments and model architectures. The source code is available under a permissive license at: github.com/linkedin/Liger-Kernel.
title Liger Kernel: Efficient Triton Kernels for LLM Training
topic Machine Learning
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2410.10989