CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiaoya, Sun, Xiaofei, Wang, Albert, Li, Jiwei, Shum, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
Tutoring LLM into a Better CUDA Optimizer
von: Brabec, Matyáš, et al.
Veröffentlicht: (2025)
von: Brabec, Matyáš, et al.
Veröffentlicht: (2025)
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
von: Su, Songqiao, et al.
Veröffentlicht: (2025)
von: Su, Songqiao, et al.
Veröffentlicht: (2025)
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2026)
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2026)
cuConv: A CUDA Implementation of Convolution for CNN Inference
von: Jordà, Marc, et al.
Veröffentlicht: (2021)
von: Jordà, Marc, et al.
Veröffentlicht: (2021)
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control
von: Buo, Antoonio, et al.
Veröffentlicht: (2026)
von: Buo, Antoonio, et al.
Veröffentlicht: (2026)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
Research on Edge Computing and Cloud Collaborative Resource Scheduling Optimization Based on Deep Reinforcement Learning
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Multimodal Federated Learning with Missing Modality via Prototype Mask and Contrast
von: Bao, Guangyin, et al.
Veröffentlicht: (2023)
von: Bao, Guangyin, et al.
Veröffentlicht: (2023)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
Beyond Aggregation: Guiding Clients in Heterogeneous Federated Learning
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning for Optimizing Energy Consumption in Smart Grid Systems
von: Alsheikhi, Abeer, et al.
Veröffentlicht: (2026)
von: Alsheikhi, Abeer, et al.
Veröffentlicht: (2026)
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
von: Li, Boyang, et al.
Veröffentlicht: (2024)
von: Li, Boyang, et al.
Veröffentlicht: (2024)
Intelligent Resource Allocation Optimization for Cloud Computing via Machine Learning
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
von: Xu, Bin, et al.
Veröffentlicht: (2026)
von: Xu, Bin, et al.
Veröffentlicht: (2026)
RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
von: Yu, Chao, et al.
Veröffentlicht: (2025)
von: Yu, Chao, et al.
Veröffentlicht: (2025)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
von: Liu, Zhihong, et al.
Veröffentlicht: (2024)
von: Liu, Zhihong, et al.
Veröffentlicht: (2024)
Enhancing Kubernetes Automated Scheduling with Deep Learning and Reinforcement Techniques for Large-Scale Cloud Computing Optimization
von: Xu, Zheng, et al.
Veröffentlicht: (2024)
von: Xu, Zheng, et al.
Veröffentlicht: (2024)
Global and Local Prompts Cooperation via Optimal Transport for Federated Learning
von: Li, Hongxia, et al.
Veröffentlicht: (2024)
von: Li, Hongxia, et al.
Veröffentlicht: (2024)
Online Parallel Multi-Task Relationship Learning via Alternating Direction Method of Multipliers
von: Li, Ruiyu, et al.
Veröffentlicht: (2024)
von: Li, Ruiyu, et al.
Veröffentlicht: (2024)
Parallel Gaussian process with kernel approximation in CUDA
von: Carminati, Davide
Veröffentlicht: (2024)
von: Carminati, Davide
Veröffentlicht: (2024)
Learn How to Query from Unlabeled Data Streams in Federated Learning
von: Sun, Yuchang, et al.
Veröffentlicht: (2024)
von: Sun, Yuchang, et al.
Veröffentlicht: (2024)
FedCGD: Collective Gradient Divergence Optimized Scheduling for Wireless Federated Learning
von: Chen, Tan, et al.
Veröffentlicht: (2025)
von: Chen, Tan, et al.
Veröffentlicht: (2025)
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
von: Mei, Zhiyu, et al.
Veröffentlicht: (2023)
von: Mei, Zhiyu, et al.
Veröffentlicht: (2023)
ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning
von: Xiao, Bangjun, et al.
Veröffentlicht: (2026)
von: Xiao, Bangjun, et al.
Veröffentlicht: (2026)
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Gradient Correction in Federated Learning with Adaptive Optimization
von: Chen, Evan, et al.
Veröffentlicht: (2025)
von: Chen, Evan, et al.
Veröffentlicht: (2025)
Loss- and Reward-Weighting for Efficient Distributed Reinforcement Learning
von: Holen, Martin, et al.
Veröffentlicht: (2023)
von: Holen, Martin, et al.
Veröffentlicht: (2023)
Deep Reinforcement Learning for System-on-Chip: Myths and Realities
von: Sung, Tegg Taekyong, et al.
Veröffentlicht: (2022)
von: Sung, Tegg Taekyong, et al.
Veröffentlicht: (2022)
A Fast and Flat Federated Learning Method via Weighted Momentum and Sharpness-Aware Minimization
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Deploying Atmospheric and Oceanic AI Models on Chinese Hardware and Framework: Migration Strategies, Performance Optimization and Analysis
von: Sun, Yuze, et al.
Veröffentlicht: (2025)
von: Sun, Yuze, et al.
Veröffentlicht: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
FedAA: A Reinforcement Learning Perspective on Adaptive Aggregation for Fair and Robust Federated Learning
von: He, Jialuo, et al.
Veröffentlicht: (2024)
von: He, Jialuo, et al.
Veröffentlicht: (2024)
Variational Bayes for Federated Continual Learning
von: Yao, Dezhong, et al.
Veröffentlicht: (2024)
von: Yao, Dezhong, et al.
Veröffentlicht: (2024)
FedImpro: Measuring and Improving Client Update in Federated Learning
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
von: Dong, Hang, et al.
Veröffentlicht: (2024)
von: Dong, Hang, et al.
Veröffentlicht: (2024)
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
von: Zhang, Zijian, et al.
Veröffentlicht: (2025) -
Tutoring LLM into a Better CUDA Optimizer
von: Brabec, Matyáš, et al.
Veröffentlicht: (2025) -
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
von: Su, Songqiao, et al.
Veröffentlicht: (2025) -
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2026) -
cuConv: A CUDA Implementation of Convolution for CNN Inference
von: Jordà, Marc, et al.
Veröffentlicht: (2021)