Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Eleftherakis, Panagiotis-Eleftherios, Anagnostopoulos, George, Kapetanakis, Anastassis, Umair, Mohammad, Vet, Jean-Yves, Iliakis, Konstantinos, Vincent, Jonathan, Gong, Jing, Patil, Akshay, García-Sánchez, Clara, Zampino, Gerardo, Vinuesa, Ricardo, Xydis, Sotirios |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dataflow Optimized Reconfigurable Acceleration for FEM-based CFD Simulations
von: Kapetanakis, Anastassis, et al.
Veröffentlicht: (2024)
von: Kapetanakis, Anastassis, et al.
Veröffentlicht: (2024)
A Bespoke Design Approach to Low-Power Printed Microprocessors for Machine Learning Applications
von: Chaidos, Panagiotis, et al.
Veröffentlicht: (2025)
von: Chaidos, Panagiotis, et al.
Veröffentlicht: (2025)
Mixed-precision Neural Networks on RISC-V Cores: ISA extensions for Multi-Pumped Soft SIMD Operations
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2024)
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2024)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
MaRVIn: A Cross-Layer Mixed-Precision RISC-V Framework for DNN Inference, from ISA Extension to Hardware Acceleration
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2025)
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2025)
Decoupled Access-Execute enabled DVFS for tinyML deployments on STM32 microcontrollers
von: Alvanaki, Elisavet Lydia, et al.
Veröffentlicht: (2024)
von: Alvanaki, Elisavet Lydia, et al.
Veröffentlicht: (2024)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
von: Alexandris, Georgios, et al.
Veröffentlicht: (2025)
von: Alexandris, Georgios, et al.
Veröffentlicht: (2025)
Multi-Partner Project: COIN-3D -- Collaborative Innovation in 3D VLSI Reliability
von: Gourdoumanis, George Rafael, et al.
Veröffentlicht: (2026)
von: Gourdoumanis, George Rafael, et al.
Veröffentlicht: (2026)
GPU Performance Portability needs Autotuning
von: Ringlein, Burkhard, et al.
Veröffentlicht: (2025)
von: Ringlein, Burkhard, et al.
Veröffentlicht: (2025)
Late Breaking Results: Leveraging Approximate Computing for Carbon-Aware DNN Accelerators
von: Panteleaki, Aikaterini Maria, et al.
Veröffentlicht: (2025)
von: Panteleaki, Aikaterini Maria, et al.
Veröffentlicht: (2025)
Carbon-Efficient 3D DNN Acceleration: Optimizing Performance and Sustainability
von: Panteleaki, Aikaterini Maria, et al.
Veröffentlicht: (2025)
von: Panteleaki, Aikaterini Maria, et al.
Veröffentlicht: (2025)
Portable Targeted Sampling Framework Using LLVM
von: Qiu, Zhantong, et al.
Veröffentlicht: (2025)
von: Qiu, Zhantong, et al.
Veröffentlicht: (2025)
Hardware Acceleration in Portable MRIs: State of the Art and Future Prospects
von: Habsi, Omar Al, et al.
Veröffentlicht: (2025)
von: Habsi, Omar Al, et al.
Veröffentlicht: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
RoboGPU: Accelerating GPU Collision Detection for Robotics
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
Analyzing Modern NVIDIA GPU cores
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Limited Read-Write/Set Hardware Transactional Memory without modifying the ISA or the Coherence Protocol
von: Kafousis, Konstantinos
Veröffentlicht: (2025)
von: Kafousis, Konstantinos
Veröffentlicht: (2025)
Design of a GPU with Heterogeneous Cores for Graphics
von: Tomás, Aurora, et al.
Veröffentlicht: (2026)
von: Tomás, Aurora, et al.
Veröffentlicht: (2026)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
von: Luo, Weile, et al.
Veröffentlicht: (2024)
von: Luo, Weile, et al.
Veröffentlicht: (2024)
COOK Access Control on an embedded Volta GPU
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
von: Fatima, Amel, et al.
Veröffentlicht: (2026)
von: Fatima, Amel, et al.
Veröffentlicht: (2026)
Multiport Support for Vortex OpenGPU Memory Hierarchy
von: Shin, Injae, et al.
Veröffentlicht: (2025)
von: Shin, Injae, et al.
Veröffentlicht: (2025)
CuLifter: Lifting GPU Binaries to Typed IR
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
always_comm: An FPGA-based Hardware Accelerator for Audio/Video Compression and Transmission
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
von: Nagendra, Savinay
Veröffentlicht: (2024)
von: Nagendra, Savinay
Veröffentlicht: (2024)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
WebAssembly on Resource-Constrained IoT Devices: Performance, Efficiency, and Portability
von: Has, Mislav, et al.
Veröffentlicht: (2025)
von: Has, Mislav, et al.
Veröffentlicht: (2025)
WebGPU-SPY: Finding Fingerprints in the Sandbox through GPU Cache Attacks
von: Ferguson, Ethan, et al.
Veröffentlicht: (2024)
von: Ferguson, Ethan, et al.
Veröffentlicht: (2024)
Leveraging Highly Approximated Multipliers in DNN Inference
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
von: Torres, L. A., et al.
Veröffentlicht: (2024)
von: Torres, L. A., et al.
Veröffentlicht: (2024)
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
von: Latif, Imran, et al.
Veröffentlicht: (2024)
von: Latif, Imran, et al.
Veröffentlicht: (2024)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dataflow Optimized Reconfigurable Acceleration for FEM-based CFD Simulations
von: Kapetanakis, Anastassis, et al.
Veröffentlicht: (2024) -
A Bespoke Design Approach to Low-Power Printed Microprocessors for Machine Learning Applications
von: Chaidos, Panagiotis, et al.
Veröffentlicht: (2025) -
Mixed-precision Neural Networks on RISC-V Cores: ISA extensions for Multi-Pumped Soft SIMD Operations
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2024) -
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
von: Leon, Vasileios, et al.
Veröffentlicht: (2025) -
MaRVIn: A Cross-Layer Mixed-Precision RISC-V Framework for DNN Inference, from ISA Extension to Hardware Acceleration
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2025)