Tutoring LLM into a Better CUDA Optimizer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Brabec, Matyáš, Klepl, Jiří, Töpfer, Michal, Kruliš, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
von: Klepl, Jiří, et al.
Veröffentlicht: (2025)
von: Klepl, Jiří, et al.
Veröffentlicht: (2025)
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control
von: Buo, Antoonio, et al.
Veröffentlicht: (2026)
von: Buo, Antoonio, et al.
Veröffentlicht: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
von: Santavas, Nicholas, et al.
Veröffentlicht: (2026)
von: Santavas, Nicholas, et al.
Veröffentlicht: (2026)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
von: Zhou, Zijie
Veröffentlicht: (2026)
von: Zhou, Zijie
Veröffentlicht: (2026)
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
von: Xu, Bin, et al.
Veröffentlicht: (2026)
von: Xu, Bin, et al.
Veröffentlicht: (2026)
xLLM Technical Report
von: Liu, Tongxuan, et al.
Veröffentlicht: (2025)
von: Liu, Tongxuan, et al.
Veröffentlicht: (2025)
Elastic On-Device LLM Service
von: Yin, Wangsong, et al.
Veröffentlicht: (2024)
von: Yin, Wangsong, et al.
Veröffentlicht: (2024)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
AI Benchmarks and Datasets for LLM Evaluation
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
High-Throughput LLM inference on Heterogeneous Clusters
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
Byzantine-Robust Decentralized Coordination of LLM Agents
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
Revisiting Parameter Server in LLM Post-Training
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
Seesaw: High-throughput LLM Inference via Model Re-sharding
von: Su, Qidong, et al.
Veröffentlicht: (2025)
von: Su, Qidong, et al.
Veröffentlicht: (2025)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
von: Jadhav, Prachi, et al.
Veröffentlicht: (2025)
von: Jadhav, Prachi, et al.
Veröffentlicht: (2025)
HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
von: Klepl, Jiří, et al.
Veröffentlicht: (2025) -
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
von: Li, Xiaoya, et al.
Veröffentlicht: (2025) -
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025) -
CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control
von: Buo, Antoonio, et al.
Veröffentlicht: (2026) -
Debunking the CUDA Myth Towards GPU-based AI Systems
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)