Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hou, Bohan, Jin, Hongyi, Wang, Guanjie, Chen, Jinqi, Cai, Yaxing, Yang, Lijie, Ye, Zihao, Ding, Yaoyao, Lai, Ruihang, Chen, Tianqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
von: Jin, Hongyi, et al.
Veröffentlicht: (2026)
von: Jin, Hongyi, et al.
Veröffentlicht: (2026)
A System for Microserving of LLMs
von: Jin, Hongyi, et al.
Veröffentlicht: (2024)
von: Jin, Hongyi, et al.
Veröffentlicht: (2024)
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
von: Klepl, Jiří, et al.
Veröffentlicht: (2025)
von: Klepl, Jiří, et al.
Veröffentlicht: (2025)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
ECDQC: Efficient Compilation for Distributed Quantum Computing with Linear Layout
von: Liu, Kecheng, et al.
Veröffentlicht: (2024)
von: Liu, Kecheng, et al.
Veröffentlicht: (2024)
UMDAM: A Unified Data Layout and DRAM Address Mapping for Heterogenous NPU-PIM
von: Huang, Hai
Veröffentlicht: (2025)
von: Huang, Hai
Veröffentlicht: (2025)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
von: Zheng, Size, et al.
Veröffentlicht: (2025)
von: Zheng, Size, et al.
Veröffentlicht: (2025)
EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
von: Yuan, Yitao, et al.
Veröffentlicht: (2026)
von: Yuan, Yitao, et al.
Veröffentlicht: (2026)
SFVInt: Simple, Fast and Generic Variable-Length Integer Decoding using Bit Manipulation Instructions
von: Liao, Gang, et al.
Veröffentlicht: (2024)
von: Liao, Gang, et al.
Veröffentlicht: (2024)
Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Serverless Abstractions for Short-Running, Lightweight Streams
von: Carl, Natalie, et al.
Veröffentlicht: (2026)
von: Carl, Natalie, et al.
Veröffentlicht: (2026)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
von: Liu, Hangda, et al.
Veröffentlicht: (2025)
von: Liu, Hangda, et al.
Veröffentlicht: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
Object Abstraction To Streamline Edge-Cloud-Native Application Development
von: Lertpongrujikorn, Pawissanutt
Veröffentlicht: (2025)
von: Lertpongrujikorn, Pawissanutt
Veröffentlicht: (2025)
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
von: Rao, Yizhuo, et al.
Veröffentlicht: (2026)
von: Rao, Yizhuo, et al.
Veröffentlicht: (2026)
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud
von: Ding, Eric
Veröffentlicht: (2025)
von: Ding, Eric
Veröffentlicht: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
Scalable Readability Evaluation for Graph Layouts: 2D Geometric Distributed Algorithms
von: Yun, Sanggeon
Veröffentlicht: (2024)
von: Yun, Sanggeon
Veröffentlicht: (2024)
WindVE: Collaborative CPU-NPU Vector Embedding
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
Efficient Parallel Compilation and Profiling of Quantum Circuits at Large Scales
von: Moore, Jane, et al.
Veröffentlicht: (2026)
von: Moore, Jane, et al.
Veröffentlicht: (2026)
LAPIS: A Performance Portable, High Productivity Compiler Framework
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
New Wide Locally Recoverable Codes with Unified Locality
von: Xu, Liangliang, et al.
Veröffentlicht: (2025)
von: Xu, Liangliang, et al.
Veröffentlicht: (2025)
Reputation-Based Leader Election under Partial Synchrony: Towards a Protocol-Independent Abstraction with Enhanced Guarantees
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
von: Lee, Sunjung, et al.
Veröffentlicht: (2026)
von: Lee, Sunjung, et al.
Veröffentlicht: (2026)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
von: Deshmukh, Aadesh, et al.
Veröffentlicht: (2025)
von: Deshmukh, Aadesh, et al.
Veröffentlicht: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
von: Yoo, Jinsun, et al.
Veröffentlicht: (2026)
von: Yoo, Jinsun, et al.
Veröffentlicht: (2026)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
FeedSign: Robust Full-parameter Federated Fine-tuning of Large Models with Extremely Low Communication Overhead of One Bit
von: Cai, Zhijie, et al.
Veröffentlicht: (2025)
von: Cai, Zhijie, et al.
Veröffentlicht: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
von: Jin, Hongyi, et al.
Veröffentlicht: (2026) -
A System for Microserving of LLMs
von: Jin, Hongyi, et al.
Veröffentlicht: (2024) -
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
von: Klepl, Jiří, et al.
Veröffentlicht: (2025) -
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025) -
ECDQC: Efficient Compilation for Distributed Quantum Computing with Linear Layout
von: Liu, Kecheng, et al.
Veröffentlicht: (2024)