Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Size, Bao, Wenlei, Hou, Qi, Zheng, Xuegui, Fang, Jin, Huang, Chenhui, Li, Tianqi, Duanmu, Haojie, Chen, Renze, Xu, Ruifan, Guo, Yifan, Zheng, Ningxin, Jiang, Ziheng, Di, Xinyi, Wang, Dongyang, Ye, Jianxi, Lin, Haibin, Chang, Li-Wen, Lu, Liqiang, Liang, Yun, Zhai, Jidong, Liu, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
by: Zheng, Size, et al.
Published: (2025)
by: Zheng, Size, et al.
Published: (2025)
DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
by: Zheng, Size, et al.
Published: (2026)
by: Zheng, Size, et al.
Published: (2026)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
by: Zheng, Size, et al.
Published: (2026)
by: Zheng, Size, et al.
Published: (2026)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
by: Li, Haonan, et al.
Published: (2025)
by: Li, Haonan, et al.
Published: (2025)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
by: Zhang, Shulai, et al.
Published: (2025)
by: Zhang, Shulai, et al.
Published: (2025)
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
by: Li, Jianling, et al.
Published: (2025)
by: Li, Jianling, et al.
Published: (2025)
ML-Triton, A Multi-Level Compilation and Language Extension to Triton GPU Programming
by: Wang, Dewei, et al.
Published: (2025)
by: Wang, Dewei, et al.
Published: (2025)
AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
by: Li, Shangzhan, et al.
Published: (2025)
by: Li, Shangzhan, et al.
Published: (2025)
The Anatomy of a Triton Attention Kernel
by: Ringlein, Burkhard, et al.
Published: (2025)
by: Ringlein, Burkhard, et al.
Published: (2025)
Liger Kernel: Efficient Triton Kernels for LLM Training
by: Hsu, Pin-Lun, et al.
Published: (2024)
by: Hsu, Pin-Lun, et al.
Published: (2024)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
by: Chang, Li-Wen, et al.
Published: (2024)
by: Chang, Li-Wen, et al.
Published: (2024)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
by: Wang, Jianghui, et al.
Published: (2025)
by: Wang, Jianghui, et al.
Published: (2025)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
by: Woo, Jiin, et al.
Published: (2025)
by: Woo, Jiin, et al.
Published: (2025)
The Humanities at Triton College.
by: Jacot, Robert E., et al.
Published: (1984)
by: Jacot, Robert E., et al.
Published: (1984)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
by: Zheng, Size, et al.
Published: (2024)
by: Zheng, Size, et al.
Published: (2024)
NineToothed: A Triton-Based High-Level Domain-Specific Language for Machine Learning
by: Huang, Jiacheng, et al.
Published: (2025)
by: Huang, Jiacheng, et al.
Published: (2025)
Sparton: Fast and Memory-Efficient Triton Kernel for Learned Sparse Retrieval
by: Nguyen, Thong, et al.
Published: (2026)
by: Nguyen, Thong, et al.
Published: (2026)
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
by: Duanmu, Haojie, et al.
Published: (2025)
by: Duanmu, Haojie, et al.
Published: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
Clouds and Hazes in the Atmospheres of Triton and Pluto
by: Gao, Peter, et al.
Published: (2024)
by: Gao, Peter, et al.
Published: (2024)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
DRTriton: Large-Scale Synthetic Data Driven Reinforcement Learning for Triton Kernel Generation
by: Guo, Siqi, et al.
Published: (2026)
by: Guo, Siqi, et al.
Published: (2026)
The trans- and post-capture orbital evolution of Triton
by: van Woerkom, Quirijn Benjamin
Published: (2024)
by: van Woerkom, Quirijn Benjamin
Published: (2024)
Triton and Pluto: same origin but separated at birth
by: Mousis, Olivier, et al.
Published: (2024)
by: Mousis, Olivier, et al.
Published: (2024)
Fast and Simplex: 2-Simplicial Attention in Triton
by: Roy, Aurko, et al.
Published: (2025)
by: Roy, Aurko, et al.
Published: (2025)
Neptune's obliquity was likely engendered by Triton's tidal evolution
by: Gomes, Rodney
Published: (2026)
by: Gomes, Rodney
Published: (2026)
Impact of the transport of magnetospheric electrons on the composition of the Triton atmosphere
by: Benne, B., et al.
Published: (2024)
by: Benne, B., et al.
Published: (2024)
New constraints on Triton's atmosphere from the 6 October 2022 stellar occultation
by: Yuan, Ye, et al.
Published: (2024)
by: Yuan, Ye, et al.
Published: (2024)
Putto With Three Tritons at The Fine Arts Museum in Brussels, Belgium
by: Scan-the-World
Published: (2026)
by: Scan-the-World
Published: (2026)
Two-Body Triton Photodisintegration and Wigner-SU(4) Symmetry
by: Lin, Xincheng, et al.
Published: (2024)
by: Lin, Xincheng, et al.
Published: (2024)
TritonDFT: Automating DFT with a Multi-Agent Framework
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Iris: First-Class Multi-GPU Programming Experience in Triton
by: Awad, Muhammad, et al.
Published: (2025)
by: Awad, Muhammad, et al.
Published: (2025)
Constraints on Triton atmospheric evolution from occultations: 1989-2022
by: Sicardy, B., et al.
Published: (2024)
by: Sicardy, B., et al.
Published: (2024)
R7 Torus Rhumb-Line Constant Part II - Cut-and-Project CSR(0) and the Tritone Kernel
by: Kirjonen, Miikka, et al.
Published: (2025)
by: Kirjonen, Miikka, et al.
Published: (2025)
Systematic Development of a Detergent Toolbox as an Alternative to Triton X‐100
by: Varsha Yadav, et al.
Published: (2025)
by: Varsha Yadav, et al.
Published: (2025)
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
by: Hoque, Adnan, et al.
Published: (2024)
by: Hoque, Adnan, et al.
Published: (2024)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
by: Jin, Chao, et al.
Published: (2025)
by: Jin, Chao, et al.
Published: (2025)
Prasugrel: la caracola de Tritón para calmar la reactividad plaquetaria
by: Antonio Tello-Montoliu
Published: (2012)
by: Antonio Tello-Montoliu
Published: (2012)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
by: Li, Cong, et al.
Published: (2025)
by: Li, Cong, et al.
Published: (2025)
Similar Items
-
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
by: Zheng, Size, et al.
Published: (2025) -
DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
by: Zheng, Size, et al.
Published: (2026) -
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
by: Zheng, Size, et al.
Published: (2026) -
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026) -
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
by: Li, Haonan, et al.
Published: (2025)