Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Yifan, Johnson, Egan, Chatarasi, Prasanth, Adve, Vikram, Misailovic, Sasa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels
por: Zhao, Yifan, et al.
Publicado: (2026)
por: Zhao, Yifan, et al.
Publicado: (2026)
RefineStat: Efficient Exploration for Probabilistic Program Synthesis
por: Kanda, Madhav, et al.
Publicado: (2025)
por: Kanda, Madhav, et al.
Publicado: (2025)
CRANE: Reasoning with constrained LLM generation
por: Banerjee, Debangshu, et al.
Publicado: (2025)
por: Banerjee, Debangshu, et al.
Publicado: (2025)
DINGO: Constrained Inference for Diffusion LLMs
por: Suresh, Tarun, et al.
Publicado: (2025)
por: Suresh, Tarun, et al.
Publicado: (2025)
IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking
por: Ugare, Shubham, et al.
Publicado: (2024)
por: Ugare, Shubham, et al.
Publicado: (2024)
Incremental Randomized Smoothing Certification
por: Ugare, Shubham, et al.
Publicado: (2023)
por: Ugare, Shubham, et al.
Publicado: (2023)
SynCode: LLM Generation with Grammar Augmentation
por: Ugare, Shubham, et al.
Publicado: (2024)
por: Ugare, Shubham, et al.
Publicado: (2024)
LEWIS (LayEr WIse Sparsity) -- A Training Free Guided Model Merging Approach
por: Chopra, Hetarth, et al.
Publicado: (2025)
por: Chopra, Hetarth, et al.
Publicado: (2025)
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
por: Hu, Muyan, et al.
Publicado: (2026)
por: Hu, Muyan, et al.
Publicado: (2026)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
por: Soi, Rupanshu, et al.
Publicado: (2025)
por: Soi, Rupanshu, et al.
Publicado: (2025)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
por: Chen, Hongzheng, et al.
Publicado: (2025)
por: Chen, Hongzheng, et al.
Publicado: (2025)
The Next 700 ML-Enabled Compiler Optimizations
por: VenkataKeerthy, S., et al.
Publicado: (2023)
por: VenkataKeerthy, S., et al.
Publicado: (2023)
Kernel Contracts: A Specification Language for ML Kernel Correctness Across Heterogeneous Silicon
por: Veit, Cooper
Publicado: (2026)
por: Veit, Cooper
Publicado: (2026)
Small Language Models as Compiler Experts: Auto-Parallelization for Heterogeneous Systems
por: Devadiga, Prathamesh
Publicado: (2025)
por: Devadiga, Prathamesh
Publicado: (2025)
Verified Lifting of Deep learning Operators
por: Zhan, Qi, et al.
Publicado: (2024)
por: Zhan, Qi, et al.
Publicado: (2024)
ARQ: A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs
por: Yang, Yuchen, et al.
Publicado: (2024)
por: Yang, Yuchen, et al.
Publicado: (2024)
Mobiprox: Supporting Dynamic Approximate Computing on Mobiles
por: Fabjančič, Matevž, et al.
Publicado: (2023)
por: Fabjančič, Matevž, et al.
Publicado: (2023)
An LLM-Tool Compiler for Fused Parallel Function Calling
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
por: Sohn, Gina, et al.
Publicado: (2025)
por: Sohn, Gina, et al.
Publicado: (2025)
A Deep Dive into Function Inlining and its Security Implications for ML-based Binary Analysis
por: Abusabha, Omar, et al.
Publicado: (2025)
por: Abusabha, Omar, et al.
Publicado: (2025)
Is The Watermarking Of LLM-Generated Code Robust?
por: Suresh, Tarun, et al.
Publicado: (2024)
por: Suresh, Tarun, et al.
Publicado: (2024)
Enforcing Temporal Constraints for LLM Agents
por: Kamath, Adharsh, et al.
Publicado: (2025)
por: Kamath, Adharsh, et al.
Publicado: (2025)
Hercules: A Compiler for Productive Programming of Heterogeneous Systems
por: Arbore, Russel, et al.
Publicado: (2025)
por: Arbore, Russel, et al.
Publicado: (2025)
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
por: Dong, Shouyang, et al.
Publicado: (2025)
por: Dong, Shouyang, et al.
Publicado: (2025)
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
por: Lacouture, Rubens, et al.
Publicado: (2025)
por: Lacouture, Rubens, et al.
Publicado: (2025)
Adaptive Branch-and-Bound Tree Exploration for Neural Network Verification
por: Fukuda, Kota, et al.
Publicado: (2025)
por: Fukuda, Kota, et al.
Publicado: (2025)
Scaling Deep Learning Training with MPMD Pipeline Parallelism
por: Xhebraj, Anxhelo, et al.
Publicado: (2024)
por: Xhebraj, Anxhelo, et al.
Publicado: (2024)
DeepCircuitX: A Comprehensive Repository-Level Dataset for RTL Code Understanding, Generation, and PPA Analysis
por: Li, Zeju, et al.
Publicado: (2025)
por: Li, Zeju, et al.
Publicado: (2025)
Worst-Case Convergence Time of ML Algorithms via Extreme Value Theory
por: Tizpaz-Niari, Saeid, et al.
Publicado: (2024)
por: Tizpaz-Niari, Saeid, et al.
Publicado: (2024)
Multi-modal Learning for WebAssembly Reverse Engineering
por: Huang, Hanxian, et al.
Publicado: (2024)
por: Huang, Hanxian, et al.
Publicado: (2024)
A Deep Learning Model for Predicting Transformation Legality
por: Tiwari, Avani, et al.
Publicado: (2025)
por: Tiwari, Avani, et al.
Publicado: (2025)
Rulebook: bringing co-routines to reinforcement learning environments
por: Fioravanti, Massimo, et al.
Publicado: (2025)
por: Fioravanti, Massimo, et al.
Publicado: (2025)
A Formally Verified Robustness Certifier for Neural Networks (Extended Version)
por: Tobler, James, et al.
Publicado: (2025)
por: Tobler, James, et al.
Publicado: (2025)
Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment
por: Boruch-Gruszecki, Aleksander, et al.
Publicado: (2025)
por: Boruch-Gruszecki, Aleksander, et al.
Publicado: (2025)
Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding
por: Guan, Yue, et al.
Publicado: (2025)
por: Guan, Yue, et al.
Publicado: (2025)
NaN-Propagation: A Novel Method for Sparsity Detection in Black-Box Computational Functions
por: Sharpe, Peter
Publicado: (2025)
por: Sharpe, Peter
Publicado: (2025)
Type-Constrained Code Generation with Language Models
por: Mündler, Niels, et al.
Publicado: (2025)
por: Mündler, Niels, et al.
Publicado: (2025)
Compiling to recurrent neurons
por: Velez-Ginorio, Joey, et al.
Publicado: (2025)
por: Velez-Ginorio, Joey, et al.
Publicado: (2025)
Guarding the Privacy of Label-Only Access to Neural Network Classifiers via iDP Verification
por: Kabaha, Anan, et al.
Publicado: (2025)
por: Kabaha, Anan, et al.
Publicado: (2025)
Verifying Computational Graphs in Production-Grade Distributed Machine Learning Frameworks
por: Zulkifli, Kahfi S., et al.
Publicado: (2025)
por: Zulkifli, Kahfi S., et al.
Publicado: (2025)
Ejemplares similares
-
Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels
por: Zhao, Yifan, et al.
Publicado: (2026) -
RefineStat: Efficient Exploration for Probabilistic Program Synthesis
por: Kanda, Madhav, et al.
Publicado: (2025) -
CRANE: Reasoning with constrained LLM generation
por: Banerjee, Debangshu, et al.
Publicado: (2025) -
DINGO: Constrained Inference for Diffusion LLMs
por: Suresh, Tarun, et al.
Publicado: (2025) -
IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking
por: Ugare, Shubham, et al.
Publicado: (2024)