On Similarity of Computational Kernels in our Codes and Proxies
Fuente:
arXiv
Saved in:
| Main Authors: | McKinsey, Michael, Brink, Stephanie, Pearce, Olga |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos
by: Nansamba, Grace, et al.
Published: (2025)
by: Nansamba, Grace, et al.
Published: (2025)
Comparing Cross-Platform Performance via Node-to-Node Scaling Studies
by: Weiss, Kenneth, et al.
Published: (2025)
by: Weiss, Kenneth, et al.
Published: (2025)
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
QEdgeProxy: QoS-Aware Load Balancing for IoT Services in the Computing Continuum
by: Čilić, Ivan, et al.
Published: (2024)
by: Čilić, Ivan, et al.
Published: (2024)
Persistent and Partitioned MPI for Stencil Communication
by: Collom, Gerald, et al.
Published: (2025)
by: Collom, Gerald, et al.
Published: (2025)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
by: Matsumura, Kazuaki, et al.
Published: (2023)
by: Matsumura, Kazuaki, et al.
Published: (2023)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
by: Andersson, Måns I., et al.
Published: (2025)
by: Andersson, Måns I., et al.
Published: (2025)
Accelerating Python Applications with Dask and ProxyStore
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
Object Proxy Patterns for Accelerating Distributed Applications
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
Code once, Run Green: Automated Green Code Translation in Serverless Computing
by: Werner, Sebastian, et al.
Published: (2025)
by: Werner, Sebastian, et al.
Published: (2025)
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
by: Zheng, Size, et al.
Published: (2025)
by: Zheng, Size, et al.
Published: (2025)
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
by: Qiu, Houming, et al.
Published: (2025)
by: Qiu, Houming, et al.
Published: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
by: Huang, Ziyu, et al.
Published: (2025)
by: Huang, Ziyu, et al.
Published: (2025)
MIDAS: Adaptive Proxy Middleware for Mitigating Metadata Hotspots in HPC I/O at Scale
by: Ghimire, Sangam, et al.
Published: (2025)
by: Ghimire, Sangam, et al.
Published: (2025)
Synthesizing Proxy Applications for MPI Programs
by: Luo, Jiyu, et al.
Published: (2023)
by: Luo, Jiyu, et al.
Published: (2023)
Privacy-Preserving Coding Schemes for Multi-Access Distributed Computing Models
by: Sasi, Shanuja
Published: (2026)
by: Sasi, Shanuja
Published: (2026)
Approximated Coded Computing: Towards Fast, Private and Secure Distributed Machine Learning
by: Qiu, Houming, et al.
Published: (2024)
by: Qiu, Houming, et al.
Published: (2024)
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
by: Zhu, Xinguo, et al.
Published: (2025)
by: Zhu, Xinguo, et al.
Published: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
by: Saurez, Enrique, et al.
Published: (2024)
by: Saurez, Enrique, et al.
Published: (2024)
Improving the Efficiency of OpenCL Kernels through Pipes
by: Zarch, Mostafa Eghbali, et al.
Published: (2022)
by: Zarch, Mostafa Eghbali, et al.
Published: (2022)
Seer: Predictive Runtime Kernel Selection for Irregular Problems
by: Swann, Ryan, et al.
Published: (2024)
by: Swann, Ryan, et al.
Published: (2024)
Measuring Data Similarity for Efficient Federated Learning: A Feasibility Study
by: Famá, Fernanda, et al.
Published: (2024)
by: Famá, Fernanda, et al.
Published: (2024)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
by: Jangda, Abhinav, et al.
Published: (2023)
by: Jangda, Abhinav, et al.
Published: (2023)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
Understanding Power and Energy Utilization in Large Scale Production Physics Simulation Codes
by: Bertsch, Adam, et al.
Published: (2022)
by: Bertsch, Adam, et al.
Published: (2022)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
by: Wahlgren, Jacob, et al.
Published: (2026)
by: Wahlgren, Jacob, et al.
Published: (2026)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024)
by: Wahlgren, Jacob, et al.
Published: (2024)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
by: Zheng, Size, et al.
Published: (2025)
by: Zheng, Size, et al.
Published: (2025)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
by: Bellavita, Julian, et al.
Published: (2025)
by: Bellavita, Julian, et al.
Published: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
Exploring Uncore Frequency Scaling for Heterogeneous Computing
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
by: Sun, Guanxiong, et al.
Published: (2025)
by: Sun, Guanxiong, et al.
Published: (2025)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
by: Zheng, Size, et al.
Published: (2026)
by: Zheng, Size, et al.
Published: (2026)
More for Less: Integrating Capability-Predominant and Capacity-Predominant Computing
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024)
by: Remke, Stefan, et al.
Published: (2024)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
by: Amoros, Oscar, et al.
Published: (2025)
by: Amoros, Oscar, et al.
Published: (2025)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024)
by: Afzal, Ayesha, et al.
Published: (2024)
Similar Items
-
Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos
by: Nansamba, Grace, et al.
Published: (2025) -
Comparing Cross-Platform Performance via Node-to-Node Scaling Studies
by: Weiss, Kenneth, et al.
Published: (2025) -
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024) -
QEdgeProxy: QoS-Aware Load Balancing for IoT Services in the Computing Continuum
by: Čilić, Ivan, et al.
Published: (2024) -
Persistent and Partitioned MPI for Stencil Communication
by: Collom, Gerald, et al.
Published: (2025)