Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yifan, Yang, Yuchen, Budiu, Matei, Misailovic, Sasa |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RefineStat: Efficient Exploration for Probabilistic Program Synthesis
by: Kanda, Madhav, et al.
Published: (2025)
by: Kanda, Madhav, et al.
Published: (2025)
Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
by: Zhao, Yifan, et al.
Published: (2025)
by: Zhao, Yifan, et al.
Published: (2025)
CRANE: Reasoning with constrained LLM generation
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
DINGO: Constrained Inference for Diffusion LLMs
by: Suresh, Tarun, et al.
Published: (2025)
by: Suresh, Tarun, et al.
Published: (2025)
IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking
by: Ugare, Shubham, et al.
Published: (2024)
by: Ugare, Shubham, et al.
Published: (2024)
Incremental Randomized Smoothing Certification
by: Ugare, Shubham, et al.
Published: (2023)
by: Ugare, Shubham, et al.
Published: (2023)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
by: Cheng, Xinhao, et al.
Published: (2025)
by: Cheng, Xinhao, et al.
Published: (2025)
SynCode: LLM Generation with Grammar Augmentation
by: Ugare, Shubham, et al.
Published: (2024)
by: Ugare, Shubham, et al.
Published: (2024)
Small Language Models as Compiler Experts: Auto-Parallelization for Heterogeneous Systems
by: Devadiga, Prathamesh
Published: (2025)
by: Devadiga, Prathamesh
Published: (2025)
LLM-Aided Compilation for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2024)
by: Hong, Charles, et al.
Published: (2024)
Hexcute: A Compiler Framework for Automating Layout Synthesis in GPU Programs
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
ARQ: A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
by: Jin, Hongyi, et al.
Published: (2026)
by: Jin, Hongyi, et al.
Published: (2026)
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
Compiling to recurrent neurons
by: Velez-Ginorio, Joey, et al.
Published: (2025)
by: Velez-Ginorio, Joey, et al.
Published: (2025)
Compiling to linear neurons
by: Velez-Ginorio, Joey, et al.
Published: (2025)
by: Velez-Ginorio, Joey, et al.
Published: (2025)
DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
by: Zheng, Size, et al.
Published: (2026)
by: Zheng, Size, et al.
Published: (2026)
CompilerDream: Learning a Compiler World Model for General Code Optimization
by: Deng, Chaoyi, et al.
Published: (2024)
by: Deng, Chaoyi, et al.
Published: (2024)
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
by: Hu, Muyan, et al.
Published: (2026)
by: Hu, Muyan, et al.
Published: (2026)
Transformers are Efficient Compilers, Provably
by: Zhai, Xiyu, et al.
Published: (2024)
by: Zhai, Xiyu, et al.
Published: (2024)
Compiler generated feedback for Large Language Models
by: Grubisic, Dejan, et al.
Published: (2024)
by: Grubisic, Dejan, et al.
Published: (2024)
Kernel Contracts: A Specification Language for ML Kernel Correctness Across Heterogeneous Silicon
by: Veit, Cooper
Published: (2026)
by: Veit, Cooper
Published: (2026)
Pattern Matching in AI Compilers and its Formalization (Extended Version)
by: Cutler, Joseph W., et al.
Published: (2024)
by: Cutler, Joseph W., et al.
Published: (2024)
PolyBlocks: A Compiler Infrastructure for AI Chips and Programming Frameworks
by: Bondhugula, Uday, et al.
Published: (2026)
by: Bondhugula, Uday, et al.
Published: (2026)
Offline Imitation Learning from Multiple Baselines with Applications to Compiler Optimization
by: Marinov, Teodor V., et al.
Published: (2024)
by: Marinov, Teodor V., et al.
Published: (2024)
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
by: Dong, Shouyang, et al.
Published: (2025)
by: Dong, Shouyang, et al.
Published: (2025)
The Next 700 ML-Enabled Compiler Optimizations
by: VenkataKeerthy, S., et al.
Published: (2023)
by: VenkataKeerthy, S., et al.
Published: (2023)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation
by: Ding, Yaoyao, et al.
Published: (2025)
by: Ding, Yaoyao, et al.
Published: (2025)
SPLAT: A framework for optimised GPU code-generation for SParse reguLar ATtention
by: Gupta, Ahan, et al.
Published: (2024)
by: Gupta, Ahan, et al.
Published: (2024)
Recursive Function Definitions in Static Dataflow Graphs and their Implementation in TensorFlow
by: Kostopoulou, Kelly, et al.
Published: (2024)
by: Kostopoulou, Kelly, et al.
Published: (2024)
WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language Models
by: Yang, Chenyuan, et al.
Published: (2023)
by: Yang, Chenyuan, et al.
Published: (2023)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
by: Younesian, Sharareh, et al.
Published: (2026)
by: Younesian, Sharareh, et al.
Published: (2026)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
An LLM-Tool Compiler for Fused Parallel Function Calling
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
SparseAuto: An Auto-Scheduler for Sparse Tensor Computations Using Recursive Loop Nest Restructuring
by: Dias, Adhitha, et al.
Published: (2023)
by: Dias, Adhitha, et al.
Published: (2023)
PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
by: Liu, Yiqun, et al.
Published: (2026)
by: Liu, Yiqun, et al.
Published: (2026)
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
by: Li, Youjie, et al.
Published: (2025)
by: Li, Youjie, et al.
Published: (2025)
Finding Missed Code Size Optimizations in Compilers using LLMs
by: Italiano, Davide, et al.
Published: (2024)
by: Italiano, Davide, et al.
Published: (2024)
Similar Items
-
RefineStat: Efficient Exploration for Probabilistic Program Synthesis
by: Kanda, Madhav, et al.
Published: (2025) -
Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
by: Zhao, Yifan, et al.
Published: (2025) -
CRANE: Reasoning with constrained LLM generation
by: Banerjee, Debangshu, et al.
Published: (2025) -
DINGO: Constrained Inference for Diffusion LLMs
by: Suresh, Tarun, et al.
Published: (2025) -
IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking
by: Ugare, Shubham, et al.
Published: (2024)