Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hou, Bohan, Jin, Hongyi, Wang, Guanjie, Chen, Jinqi, Cai, Yaxing, Yang, Lijie, Ye, Zihao, Ding, Yaoyao, Lai, Ruihang, Chen, Tianqi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
par: Jin, Hongyi, et autres
Publié: (2026)
par: Jin, Hongyi, et autres
Publié: (2026)
A System for Microserving of LLMs
par: Jin, Hongyi, et autres
Publié: (2024)
par: Jin, Hongyi, et autres
Publié: (2024)
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
par: Klepl, Jiří, et autres
Publié: (2025)
par: Klepl, Jiří, et autres
Publié: (2025)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
par: Cheng, Xinhao, et autres
Publié: (2025)
par: Cheng, Xinhao, et autres
Publié: (2025)
ECDQC: Efficient Compilation for Distributed Quantum Computing with Linear Layout
par: Liu, Kecheng, et autres
Publié: (2024)
par: Liu, Kecheng, et autres
Publié: (2024)
UMDAM: A Unified Data Layout and DRAM Address Mapping for Heterogenous NPU-PIM
par: Huang, Hai
Publié: (2025)
par: Huang, Hai
Publié: (2025)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
par: Zheng, Size, et autres
Publié: (2025)
par: Zheng, Size, et autres
Publié: (2025)
EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
par: Yuan, Yitao, et autres
Publié: (2026)
par: Yuan, Yitao, et autres
Publié: (2026)
SFVInt: Simple, Fast and Generic Variable-Length Integer Decoding using Bit Manipulation Instructions
par: Liao, Gang, et autres
Publié: (2024)
par: Liao, Gang, et autres
Publié: (2024)
Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey
par: Chen, Chen, et autres
Publié: (2024)
par: Chen, Chen, et autres
Publié: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
par: Bhosale, Aditya, et autres
Publié: (2025)
par: Bhosale, Aditya, et autres
Publié: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
par: Ye, Zihao, et autres
Publié: (2025)
par: Ye, Zihao, et autres
Publié: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
Serverless Abstractions for Short-Running, Lightweight Streams
par: Carl, Natalie, et autres
Publié: (2026)
par: Carl, Natalie, et autres
Publié: (2026)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
par: Huang, Lexiang, et autres
Publié: (2024)
par: Huang, Lexiang, et autres
Publié: (2024)
Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
par: Liu, Hangda, et autres
Publié: (2025)
par: Liu, Hangda, et autres
Publié: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
par: Tanaka, Masahiro, et autres
Publié: (2025)
par: Tanaka, Masahiro, et autres
Publié: (2025)
Object Abstraction To Streamline Edge-Cloud-Native Application Development
par: Lertpongrujikorn, Pawissanutt
Publié: (2025)
par: Lertpongrujikorn, Pawissanutt
Publié: (2025)
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
par: Rao, Yizhuo, et autres
Publié: (2026)
par: Rao, Yizhuo, et autres
Publié: (2026)
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
par: Chen, Zihao, et autres
Publié: (2025)
par: Chen, Zihao, et autres
Publié: (2025)
Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud
par: Ding, Eric
Publié: (2025)
par: Ding, Eric
Publié: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
par: Bridges, Patrick G., et autres
Publié: (2026)
par: Bridges, Patrick G., et autres
Publié: (2026)
Scalable Readability Evaluation for Graph Layouts: 2D Geometric Distributed Algorithms
par: Yun, Sanggeon
Publié: (2024)
par: Yun, Sanggeon
Publié: (2024)
WindVE: Collaborative CPU-NPU Vector Embedding
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
par: Zhao, Bohan, et autres
Publié: (2025)
par: Zhao, Bohan, et autres
Publié: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
par: He, Yongchao, et autres
Publié: (2025)
par: He, Yongchao, et autres
Publié: (2025)
Efficient Parallel Compilation and Profiling of Quantum Circuits at Large Scales
par: Moore, Jane, et autres
Publié: (2026)
par: Moore, Jane, et autres
Publié: (2026)
LAPIS: A Performance Portable, High Productivity Compiler Framework
par: Kelley, Brian, et autres
Publié: (2025)
par: Kelley, Brian, et autres
Publié: (2025)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
par: Zhao, Bohan, et autres
Publié: (2025)
par: Zhao, Bohan, et autres
Publié: (2025)
New Wide Locally Recoverable Codes with Unified Locality
par: Xu, Liangliang, et autres
Publié: (2025)
par: Xu, Liangliang, et autres
Publié: (2025)
Reputation-Based Leader Election under Partial Synchrony: Towards a Protocol-Independent Abstraction with Enhanced Guarantees
par: Liu, Xuyang, et autres
Publié: (2025)
par: Liu, Xuyang, et autres
Publié: (2025)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
par: Lee, Sunjung, et autres
Publié: (2026)
par: Lee, Sunjung, et autres
Publié: (2026)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
par: Deshmukh, Aadesh, et autres
Publié: (2025)
par: Deshmukh, Aadesh, et autres
Publié: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
par: Yoo, Jinsun, et autres
Publié: (2026)
par: Yoo, Jinsun, et autres
Publié: (2026)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
par: Zhao, Xinkui, et autres
Publié: (2025)
par: Zhao, Xinkui, et autres
Publié: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
par: Li, Maoliang, et autres
Publié: (2026)
par: Li, Maoliang, et autres
Publié: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
par: Lv, Jiaqi, et autres
Publié: (2025)
par: Lv, Jiaqi, et autres
Publié: (2025)
FeedSign: Robust Full-parameter Federated Fine-tuning of Large Models with Extremely Low Communication Overhead of One Bit
par: Cai, Zhijie, et autres
Publié: (2025)
par: Cai, Zhijie, et autres
Publié: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
par: Wang, Chao, et autres
Publié: (2025)
par: Wang, Chao, et autres
Publié: (2025)
Documents similaires
-
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
par: Jin, Hongyi, et autres
Publié: (2026) -
A System for Microserving of LLMs
par: Jin, Hongyi, et autres
Publié: (2024) -
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
par: Klepl, Jiří, et autres
Publié: (2025) -
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
par: Cheng, Xinhao, et autres
Publié: (2025) -
ECDQC: Efficient Compilation for Distributed Quantum Computing with Linear Layout
par: Liu, Kecheng, et autres
Publié: (2024)