Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Reinhard, Anton, Ehrig, Simeon, Widera, René, Bussmann, Michael, Acosta, Uwe Hernandez |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Extrae.jl: Julia bindings for the Extrae HPC Profiler
von: Sanchez-Ramirez, Sergio, et al.
Veröffentlicht: (2025)
von: Sanchez-Ramirez, Sergio, et al.
Veröffentlicht: (2025)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
Inductive Loop Analysis for Practical HPC Application Optimization
von: Schaad, Philipp, et al.
Veröffentlicht: (2025)
von: Schaad, Philipp, et al.
Veröffentlicht: (2025)
Enabling High-Throughput Parallel I/O in Particle-in-Cell Monte Carlo Simulations with openPMD and Darshan I/O Monitoring
von: Williams, Jeremy J., et al.
Veröffentlicht: (2024)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
Energy-Aware Computing in the Year 2026
von: Tchakoute, Roblex Nana, et al.
Veröffentlicht: (2026)
von: Tchakoute, Roblex Nana, et al.
Veröffentlicht: (2026)
Hiku: Pull-Based Scheduling for Serverless Computing
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
Synthesizing Proxy Applications for MPI Programs
von: Luo, Jiyu, et al.
Veröffentlicht: (2023)
von: Luo, Jiyu, et al.
Veröffentlicht: (2023)
Usability Evaluation of Cloud for HPC Applications
von: Sochat, Vanessa, et al.
Veröffentlicht: (2025)
von: Sochat, Vanessa, et al.
Veröffentlicht: (2025)
Optimal Configuration of API Resources in Cloud Native Computing
von: Truyen, Eddy, et al.
Veröffentlicht: (2025)
von: Truyen, Eddy, et al.
Veröffentlicht: (2025)
Cloud Resource Allocation with Convex Optimization
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
Denoising Application Performance Models with Noise-Resilient Priors
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
Universal Workers: A Vision for Eliminating Cold Starts in Serverless Computing
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
Modeling the Effect of Data Redundancy on Speedup in MLFMA Near-Field Computation
von: Sadeghi, Morteza
Veröffentlicht: (2025)
von: Sadeghi, Morteza
Veröffentlicht: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
von: Tharwani, Jay, et al.
Veröffentlicht: (2024)
von: Tharwani, Jay, et al.
Veröffentlicht: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
DREAMS: Decentralized Resource Allocation and Service Management across the Compute Continuum Using Service Affinity
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
The SAP Cloud Infrastructure Dataset: A Reality Check of Scheduling and Placement of VMs in Cloud Computing
von: Uhlig, Arno, et al.
Veröffentlicht: (2025)
von: Uhlig, Arno, et al.
Veröffentlicht: (2025)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026)
von: Chen, David, et al.
Veröffentlicht: (2026)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Extrae.jl: Julia bindings for the Extrae HPC Profiler
von: Sanchez-Ramirez, Sergio, et al.
Veröffentlicht: (2025) -
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026) -
Inductive Loop Analysis for Practical HPC Application Optimization
von: Schaad, Philipp, et al.
Veröffentlicht: (2025) -
Enabling High-Throughput Parallel I/O in Particle-in-Cell Monte Carlo Simulations with openPMD and Darshan I/O Monitoring
von: Williams, Jeremy J., et al.
Veröffentlicht: (2024) -
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)