Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Hongyi, Hou, Bohan, Wang, Guanjie, Lai, Ruihang, Chen, Jinqi, Ye, Zihao, Cai, Yaxing, Dong, Yixin, Cheng, Xinhao, Zhang, Zhihao, Zhao, Yilong, Huang, Yingyi, Yang, Lijie, Jiang, Jinchen, Oliaro, Gabriele, Ji, Jianan, Miao, Xupeng, Grover, Vinod, Mowry, Todd C., Jia, Zhihao, Chen, Tianqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
by: Hou, Bohan, et al.
Published: (2026)
by: Hou, Bohan, et al.
Published: (2026)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
by: Cheng, Xinhao, et al.
Published: (2025)
by: Cheng, Xinhao, et al.
Published: (2025)
A System for Microserving of LLMs
by: Jin, Hongyi, et al.
Published: (2024)
by: Jin, Hongyi, et al.
Published: (2024)
ACRoBat: Optimizing Auto-batching of Dynamic Deep Learning at Compile Time
by: Fegade, Pratik, et al.
Published: (2023)
by: Fegade, Pratik, et al.
Published: (2023)
Relax: Composable Abstractions for End-to-End Dynamic Machine Learning
by: Lai, Ruihang, et al.
Published: (2023)
by: Lai, Ruihang, et al.
Published: (2023)
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
by: Dong, Yixin, et al.
Published: (2024)
by: Dong, Yixin, et al.
Published: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
by: Chowdhary, Sangeeta, et al.
Published: (2026)
by: Chowdhary, Sangeeta, et al.
Published: (2026)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
by: Zhang, Zhengxin, et al.
Published: (2024)
by: Zhang, Zhengxin, et al.
Published: (2024)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
PithTrain: A Compact and Agent-Native MoE Training System
by: Lai, Ruihang, et al.
Published: (2026)
by: Lai, Ruihang, et al.
Published: (2026)
Modeling Layout Abstractions Using Integer Set Relations
by: Bhaskaracharya, Somashekaracharya G, et al.
Published: (2025)
by: Bhaskaracharya, Somashekaracharya G, et al.
Published: (2025)
Mirage: A Multi-Level Superoptimizer for Tensor Programs
by: Wu, Mengdi, et al.
Published: (2024)
by: Wu, Mengdi, et al.
Published: (2024)
Megakernel vs Wavefront GPU Path Tracing
by: Padilla, Rafael, et al.
Published: (2026)
by: Padilla, Rafael, et al.
Published: (2026)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
Solsmith: Solidity Random Program Generator for Compiler Testing
by: Li, Lantian, et al.
Published: (2025)
by: Li, Lantian, et al.
Published: (2025)
DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction
by: Tang, Hongyi, et al.
Published: (2026)
by: Tang, Hongyi, et al.
Published: (2026)
Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework
by: Tang, Hongyi, et al.
Published: (2025)
by: Tang, Hongyi, et al.
Published: (2025)
XGrammar-2: Efficient Dynamic Structured Generation Engine for Agentic LLMs
by: Li, Linzhang, et al.
Published: (2026)
by: Li, Linzhang, et al.
Published: (2026)
Pattern Matching in AI Compilers and its Formalization (Extended Version)
by: Cutler, Joseph W., et al.
Published: (2024)
by: Cutler, Joseph W., et al.
Published: (2024)
Eliminating Hidden Serialization in Multi-Node Megakernel Communication
by: Oh, Byungsoo, et al.
Published: (2026)
by: Oh, Byungsoo, et al.
Published: (2026)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
by: Yang, Lijie, et al.
Published: (2025)
by: Yang, Lijie, et al.
Published: (2025)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
by: Sharma, Vyom, et al.
Published: (2026)
by: Sharma, Vyom, et al.
Published: (2026)
A Performance Model for Warp Specialization Kernels
by: Liu, Zhengyang, et al.
Published: (2025)
by: Liu, Zhengyang, et al.
Published: (2025)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
Optimal Pricing with Unreliable Signals
by: Tang, Zhihao Gavin, et al.
Published: (2026)
by: Tang, Zhihao Gavin, et al.
Published: (2026)
Pricing with a Hidden Sample
by: Tang, Zhihao Gavin, et al.
Published: (2026)
by: Tang, Zhihao Gavin, et al.
Published: (2026)
Optimal Kernel Orchestration for Tensor Programs with Korch
by: Hu, Muyan, et al.
Published: (2024)
by: Hu, Muyan, et al.
Published: (2024)
Classifications and bifurcations of tangent points and their loops of planar piecewise-smooth systems
by: Fang, Zhihao, et al.
Published: (2024)
by: Fang, Zhihao, et al.
Published: (2024)
MambaUIE&SR: Unraveling the Ocean's Secrets with Only 2.8 GFLOPs
by: Chen, Zhihao, et al.
Published: (2024)
by: Chen, Zhihao, et al.
Published: (2024)
Part-Attention Based Model Make Occluded Person Re-Identification Stronger
by: Chen, Zhihao, et al.
Published: (2024)
by: Chen, Zhihao, et al.
Published: (2024)
Bifurcations and explicit unfoldings of grazing loops connecting one high multiplicity tangent point
by: Fang, Zhihao, et al.
Published: (2024)
by: Fang, Zhihao, et al.
Published: (2024)
RAP: Retrieve, Adapt, and Prompt-Fit for Training-Free Few-Shot Medical Image Segmentation
by: Mao, Zhihao, et al.
Published: (2026)
by: Mao, Zhihao, et al.
Published: (2026)
Accelerating Retrieval-Augmented Language Model Serving with Speculation
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
Atlas: Hierarchical Partitioning for Quantum Circuit Simulation on GPUs (Extended Version)
by: Xu, Mingkuan, et al.
Published: (2024)
by: Xu, Mingkuan, et al.
Published: (2024)
Similar Items
-
Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
by: Hou, Bohan, et al.
Published: (2026) -
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023) -
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
by: Cheng, Xinhao, et al.
Published: (2025) -
A System for Microserving of LLMs
by: Jin, Hongyi, et al.
Published: (2024) -
ACRoBat: Optimizing Auto-batching of Dynamic Deep Learning at Compile Time
by: Fegade, Pratik, et al.
Published: (2023)