Temporal Load Imbalance on Ondes3D Seismic Simulator for Different Multicore Architectures
Fuente:
arXiv
Salvato in:
| Autori principali: | Solórzano, Ana Luisa Veroneze, Navaux, Philippe Olivier Alexandre, Schnorr, Lucas Mello |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
di: Miletto, Marcelo Cogo, et al.
Pubblicazione: (2024)
di: Miletto, Marcelo Cogo, et al.
Pubblicazione: (2024)
Communication-Aware Diffusion Load Balancing for Persistently Interacting Objects
di: Taylor, Maya, et al.
Pubblicazione: (2026)
di: Taylor, Maya, et al.
Pubblicazione: (2026)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
di: Apanasevich, L., et al.
Pubblicazione: (2024)
di: Apanasevich, L., et al.
Pubblicazione: (2024)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2025)
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2025)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
di: Ramesh, Risshab Srinivas
Pubblicazione: (2024)
di: Ramesh, Risshab Srinivas
Pubblicazione: (2024)
An Experimental Study of Different Aggregation Schemes in Semi-Asynchronous Federated Learning
di: Li, Yunbo, et al.
Pubblicazione: (2024)
di: Li, Yunbo, et al.
Pubblicazione: (2024)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
di: Sakib, Nazmus, et al.
Pubblicazione: (2025)
di: Sakib, Nazmus, et al.
Pubblicazione: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
di: Suffa, Philipp, et al.
Pubblicazione: (2024)
di: Suffa, Philipp, et al.
Pubblicazione: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
di: Yildiz, Mert, et al.
Pubblicazione: (2025)
di: Yildiz, Mert, et al.
Pubblicazione: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
di: Vatsavai, Sairam Sri, et al.
Pubblicazione: (2025)
di: Vatsavai, Sairam Sri, et al.
Pubblicazione: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
di: Owen, Herbert, et al.
Pubblicazione: (2024)
di: Owen, Herbert, et al.
Pubblicazione: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
di: Checconi, Fabio, et al.
Pubblicazione: (2022)
di: Checconi, Fabio, et al.
Pubblicazione: (2022)
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
di: Singapuram, Sanjay Sri Vallabh, et al.
Pubblicazione: (2025)
di: Singapuram, Sanjay Sri Vallabh, et al.
Pubblicazione: (2025)
Kairos: Efficient Temporal Graph Analytics on a Single Machine
di: da Trindade, Joana M. F., et al.
Pubblicazione: (2024)
di: da Trindade, Joana M. F., et al.
Pubblicazione: (2024)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
di: Cornelius, Melanie, et al.
Pubblicazione: (2025)
di: Cornelius, Melanie, et al.
Pubblicazione: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
di: Liu, Shifang, et al.
Pubblicazione: (2025)
di: Liu, Shifang, et al.
Pubblicazione: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
di: Mao, Ying, et al.
Pubblicazione: (2020)
di: Mao, Ying, et al.
Pubblicazione: (2020)
Staging Blocked Evaluation over Structured Sparse Matrices
di: Das, Pratyush, et al.
Pubblicazione: (2024)
di: Das, Pratyush, et al.
Pubblicazione: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
di: Debnath, Shimul, et al.
Pubblicazione: (2026)
di: Debnath, Shimul, et al.
Pubblicazione: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
di: Ng, Nathan, et al.
Pubblicazione: (2026)
di: Ng, Nathan, et al.
Pubblicazione: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
di: Jeffery, Andrew, et al.
Pubblicazione: (2024)
di: Jeffery, Andrew, et al.
Pubblicazione: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
Bridding OT and PaaS in Edge-to-Cloud Continuum
di: Barrios, Carlos J, et al.
Pubblicazione: (2025)
di: Barrios, Carlos J, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
di: Kabamba, Herve Mbikayi, et al.
Pubblicazione: (2023)
di: Kabamba, Herve Mbikayi, et al.
Pubblicazione: (2023)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
di: Landau, Diogo, et al.
Pubblicazione: (2026)
di: Landau, Diogo, et al.
Pubblicazione: (2026)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
di: Berg, Benjamin, et al.
Pubblicazione: (2024)
di: Berg, Benjamin, et al.
Pubblicazione: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
di: Miletto, Marcelo Cogo, et al.
Pubblicazione: (2024) -
Communication-Aware Diffusion Load Balancing for Persistently Interacting Objects
di: Taylor, Maya, et al.
Pubblicazione: (2026) -
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
di: Apanasevich, L., et al.
Pubblicazione: (2024) -
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2025) -
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
di: Bogart, Christopher, et al.
Pubblicazione: (2025)