Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fatima, Amel, Ta, Tuan, Beckmann, Bradford M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
HetGPU: The pursuit of making binary compatibility towards GPUs
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
Handling of Memory Page Faults during Virtual-Address RDMA
von: Psistakis, Antonis
Veröffentlicht: (2025)
von: Psistakis, Antonis
Veröffentlicht: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
IOMMU Support for Virtual-Address Remote DMA in an ARMv8 environment
von: Psistakis, Antonis
Veröffentlicht: (2025)
von: Psistakis, Antonis
Veröffentlicht: (2025)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
von: Torres, L. A., et al.
Veröffentlicht: (2024)
von: Torres, L. A., et al.
Veröffentlicht: (2024)
Analyzing a Two-Tier Disaggregated Memory Protection Scheme Based on Memory Replication
von: Volos, Haris, et al.
Veröffentlicht: (2025)
von: Volos, Haris, et al.
Veröffentlicht: (2025)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
von: Srimani, Tathagata, et al.
Veröffentlicht: (2024)
von: Srimani, Tathagata, et al.
Veröffentlicht: (2024)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
Muchisim: A Simulation Framework for Design Exploration of Multi-Chip Manycore Systems
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
von: Wang, Tinglue, et al.
Veröffentlicht: (2025)
von: Wang, Tinglue, et al.
Veröffentlicht: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
FPGA or GPU? Analyzing comparative research for application-specific guidance
von: Purkayastha, Arnab A, et al.
Veröffentlicht: (2025)
von: Purkayastha, Arnab A, et al.
Veröffentlicht: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
Enhancing Regression Models for Complex Systems Using Evolutionary Techniques for Feature Engineering
von: Arroba, Patricia, et al.
Veröffentlicht: (2024)
von: Arroba, Patricia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023) -
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026) -
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026) -
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026) -
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
von: Li, Bingyao, et al.
Veröffentlicht: (2024)