Seer: Predictive Runtime Kernel Selection for Irregular Problems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Swann, Ryan, Osama, Muhammad, Sangaiah, Karthik, Mahmud, Jalal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
von: Swann, Ryan, et al.
Veröffentlicht: (2025)
von: Swann, Ryan, et al.
Veröffentlicht: (2025)
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
von: Trifan, Octavian Alexandru, et al.
Veröffentlicht: (2025)
von: Trifan, Octavian Alexandru, et al.
Veröffentlicht: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2024)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2024)
Fair Kernel-Lock-Free Claim/Release Protocol for Shared Object Access in Cooperatively Scheduled Runtimes
von: Chalmers, Kevin, et al.
Veröffentlicht: (2025)
von: Chalmers, Kevin, et al.
Veröffentlicht: (2025)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
von: Qin, Ruoyu, et al.
Veröffentlicht: (2025)
von: Qin, Ruoyu, et al.
Veröffentlicht: (2025)
Automatic Tracing in Task-Based Runtime Systems
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
An AI-Native Runtime for Multi-Wearable Environments
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
On the Runtime of Local Mutual Exclusion for Anonymous Dynamic Networks
von: Chaturvedi, Anya, et al.
Veröffentlicht: (2025)
von: Chaturvedi, Anya, et al.
Veröffentlicht: (2025)
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
Asynchronous Wait-Free Runtime Verification and Enforcement of Linearizability
von: Castañeda, Armando, et al.
Veröffentlicht: (2023)
von: Castañeda, Armando, et al.
Veröffentlicht: (2023)
Iris: First-Class Multi-GPU Programming Experience in Triton
von: Awad, Muhammad, et al.
Veröffentlicht: (2025)
von: Awad, Muhammad, et al.
Veröffentlicht: (2025)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
RIMMS: Runtime Integrated Memory Management System for Heterogeneous Computing
von: Gener, Serhan, et al.
Veröffentlicht: (2025)
von: Gener, Serhan, et al.
Veröffentlicht: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
A Structure-Aware Irregular Blocking Method for Sparse LU Factorization
von: Hu, Zhen, et al.
Veröffentlicht: (2025)
von: Hu, Zhen, et al.
Veröffentlicht: (2025)
Privacy-Preserving Sharing of Data Analytics Runtime Metrics for Performance Modeling
von: Will, Jonathan, et al.
Veröffentlicht: (2024)
von: Will, Jonathan, et al.
Veröffentlicht: (2024)
Asynchronous Fault-Tolerant Language Decidability for Runtime Verification of Distributed Systems
von: Castañeda, Armando, et al.
Veröffentlicht: (2025)
von: Castañeda, Armando, et al.
Veröffentlicht: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
von: Stepanek, Lukas
Veröffentlicht: (2026)
von: Stepanek, Lukas
Veröffentlicht: (2026)
WANify: Gauging and Balancing Runtime WAN Bandwidth for Geo-distributed Data Analytics
von: Mohapatra, Anshuman Das, et al.
Veröffentlicht: (2025)
von: Mohapatra, Anshuman Das, et al.
Veröffentlicht: (2025)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
von: Park, Seongyeon, et al.
Veröffentlicht: (2025)
von: Park, Seongyeon, et al.
Veröffentlicht: (2025)
Lumos: Performance Characterization of WebAssembly as a Serverless Runtime in the Edge-Cloud Continuum
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
von: Da, Wei, et al.
Veröffentlicht: (2026)
von: Da, Wei, et al.
Veröffentlicht: (2026)
Overcoming Latency-bound Limitations of Distributed Graph Algorithms using the HPX Runtime System
von: Mohammadiporshokooh, Karame, et al.
Veröffentlicht: (2026)
von: Mohammadiporshokooh, Karame, et al.
Veröffentlicht: (2026)
Radiation Hydrodynamics at Scale: Comparing MPI and Asynchronous Many-Task Runtimes with FleCSI
von: Strack, Alexander, et al.
Veröffentlicht: (2026)
von: Strack, Alexander, et al.
Veröffentlicht: (2026)
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
von: Niam, Arefin, et al.
Veröffentlicht: (2026)
von: Niam, Arefin, et al.
Veröffentlicht: (2026)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
Communication Efficient Byzantine Agreement with Predictions
von: Dzulfikar, Muhammad Ayaz, et al.
Veröffentlicht: (2026)
von: Dzulfikar, Muhammad Ayaz, et al.
Veröffentlicht: (2026)
INSPIRIT: Optimizing Heterogeneous Task Scheduling through Adaptive Priority in Task-based Runtime Systems
von: Wang, Yiqing, et al.
Veröffentlicht: (2024)
von: Wang, Yiqing, et al.
Veröffentlicht: (2024)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
von: Shan, Baodi, et al.
Veröffentlicht: (2026)
von: Shan, Baodi, et al.
Veröffentlicht: (2026)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
Channel Prediction under Network Distribution Shift Using Continual Learning-based Loss Regularization
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2025)
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2025)
Optimising Blockchain Scalability for Real-Time IoT Applications
von: Rhidoy, Hasan Mahmud, et al.
Veröffentlicht: (2026)
von: Rhidoy, Hasan Mahmud, et al.
Veröffentlicht: (2026)
CWASI: A WebAssembly Runtime Shim for Inter-function Communication in the Serverless Edge-Cloud Continuum
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
von: Swann, Ryan, et al.
Veröffentlicht: (2025) -
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
von: Trifan, Octavian Alexandru, et al.
Veröffentlicht: (2025) -
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
von: Saurez, Enrique, et al.
Veröffentlicht: (2024) -
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025) -
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2024)