NineToothed: A Triton-Based High-Level Domain-Specific Language for Machine Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jiacheng, Li, Zimin, Li, Yinghui, Wang, Haojie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
by: Zheng, Size, et al.
Published: (2025)
by: Zheng, Size, et al.
Published: (2025)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
by: Zhu, Jianian, et al.
Published: (2025)
by: Zhu, Jianian, et al.
Published: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
by: Reinhard, Anton, et al.
Published: (2025)
by: Reinhard, Anton, et al.
Published: (2025)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
by: Jia, Haojie, et al.
Published: (2025)
by: Jia, Haojie, et al.
Published: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)
by: Chen, Jiefei, et al.
Published: (2026)
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Iris: First-Class Multi-GPU Programming Experience in Triton
by: Awad, Muhammad, et al.
Published: (2025)
by: Awad, Muhammad, et al.
Published: (2025)
BlazingAML: High-Throughput Anti-Money Laundering (AML) via Multi-Stage Graph Mining
by: Ye, Haojie, et al.
Published: (2026)
by: Ye, Haojie, et al.
Published: (2026)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
by: Liu, Jinyang, et al.
Published: (2023)
by: Liu, Jinyang, et al.
Published: (2023)
Experiences Building Enterprise-Level Privacy-Preserving Federated Learning to Power AI for Science
by: Li, Zilinghan, et al.
Published: (2025)
by: Li, Zilinghan, et al.
Published: (2025)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
BFLN: A Blockchain-based Federated Learning Model for Non-IID Data
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
by: Wang, Hongyu, et al.
Published: (2026)
by: Wang, Hongyu, et al.
Published: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
by: Han, Pengchao, et al.
Published: (2025)
by: Han, Pengchao, et al.
Published: (2025)
A Simulated Annealing Approach to Identical Parallel Machine Scheduling
by: Li, Jiaxing, et al.
Published: (2024)
by: Li, Jiaxing, et al.
Published: (2024)
A Knowledge Distillation-empowered Adaptive Federated Reinforcement Learning Framework for Multi-Domain IoT Applications Scheduling
by: Wang, Zhiyu, et al.
Published: (2025)
by: Wang, Zhiyu, et al.
Published: (2025)
Mazzaroth: A High-Throughput DAG Consensus with State Root
by: Li, Haohan
Published: (2025)
by: Li, Haohan
Published: (2025)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
by: Zhang, Runhua, et al.
Published: (2025)
by: Zhang, Runhua, et al.
Published: (2025)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
by: Knorr, Fabian, et al.
Published: (2025)
by: Knorr, Fabian, et al.
Published: (2025)
WindGP: Efficient Graph Partitioning on Heterogenous Machines
by: Zeng, Li, et al.
Published: (2024)
by: Zeng, Li, et al.
Published: (2024)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
by: Li, Pengbo, et al.
Published: (2026)
by: Li, Pengbo, et al.
Published: (2026)
EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel Execution
by: Bourgeois, Daniel, et al.
Published: (2024)
by: Bourgeois, Daniel, et al.
Published: (2024)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
AMECOS: A Modular Event-based Framework for Concurrent Object Specification
by: Albouy, Timothé, et al.
Published: (2024)
by: Albouy, Timothé, et al.
Published: (2024)
SplitLLM: Hierarchical Split Learning for Large Language Model over Wireless Network
by: Zhang, Songge, et al.
Published: (2025)
by: Zhang, Songge, et al.
Published: (2025)
Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
by: Gupta, Sushant Kumar, et al.
Published: (2025)
by: Gupta, Sushant Kumar, et al.
Published: (2025)
Adaptive Cache Management for Complex Storage Systems Using CNN-LSTM-Based Spatiotemporal Prediction
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
by: Hoque, Adnan, et al.
Published: (2024)
by: Hoque, Adnan, et al.
Published: (2024)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)
by: Chen, Qiaoling, et al.
Published: (2026)
Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud
by: Ding, Eric
Published: (2025)
by: Ding, Eric
Published: (2025)
ShardTensor: Domain Parallelism for Scientific Machine Learning
by: Adams, Corey, et al.
Published: (2026)
by: Adams, Corey, et al.
Published: (2026)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
Learning Process Energy Profiles from Node-Level Power Data
by: Bader, Jonathan, et al.
Published: (2025)
by: Bader, Jonathan, et al.
Published: (2025)
A Proposal for High-Level Architectural Model Capable of Expressing Various Data Collaboration Platform and Data Space Concepts
by: Dobashi, Masaru, et al.
Published: (2025)
by: Dobashi, Masaru, et al.
Published: (2025)
A Privacy-Preserving Machine Learning Framework for Edge Intelligence: An Empirical Analysis
by: Trieu, Quoc Lap, et al.
Published: (2026)
by: Trieu, Quoc Lap, et al.
Published: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Similar Items
-
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
by: Zheng, Size, et al.
Published: (2025) -
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
by: Wu, Hang, et al.
Published: (2025) -
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
by: Zhu, Jianian, et al.
Published: (2025) -
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025) -
Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
by: Reinhard, Anton, et al.
Published: (2025)