Joint Training on AMD and NVIDIA GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Jon, Jia, Thomas, Zhu, Jing, Yu, Zhendong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
von: Tramm, John, et al.
Veröffentlicht: (2024)
von: Tramm, John, et al.
Veröffentlicht: (2024)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2026)
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2026)
Dynamic Memory Management on GPUs with SYCL
von: Standish, Russell K.
Veröffentlicht: (2025)
von: Standish, Russell K.
Veröffentlicht: (2025)
Improving inference time in multi-TPU systems with profiled model segmentation
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2025)
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2025)
Balanced segmentation of CNNs for multi-TPU inference
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2025)
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Stream parallel skeleton optimization
von: Aldinucci, Marco, et al.
Veröffentlicht: (2024)
von: Aldinucci, Marco, et al.
Veröffentlicht: (2024)
StreamFlow: cross-breeding cloud with HPC
von: Colonnelli, Iacopo, et al.
Veröffentlicht: (2020)
von: Colonnelli, Iacopo, et al.
Veröffentlicht: (2020)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
von: Li, Chendi, et al.
Veröffentlicht: (2022)
von: Li, Chendi, et al.
Veröffentlicht: (2022)
Optimizing sDTW for AMD GPUs
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
von: Wang, Wenyi, et al.
Veröffentlicht: (2025)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
von: Xu, Yao, et al.
Veröffentlicht: (2024)
von: Xu, Yao, et al.
Veröffentlicht: (2024)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
von: Gallone, Anna, et al.
Veröffentlicht: (2025)
von: Gallone, Anna, et al.
Veröffentlicht: (2025)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
von: Mantha, Pradeep, et al.
Veröffentlicht: (2026)
von: Mantha, Pradeep, et al.
Veröffentlicht: (2026)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
von: Ma, Cong, et al.
Veröffentlicht: (2025)
von: Ma, Cong, et al.
Veröffentlicht: (2025)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
von: Li, Yinghan, et al.
Veröffentlicht: (2025)
von: Li, Yinghan, et al.
Veröffentlicht: (2025)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
Scalable Concurrent Queues for GPU
von: Shetty, Pratheek Prakash, et al.
Veröffentlicht: (2026)
von: Shetty, Pratheek Prakash, et al.
Veröffentlicht: (2026)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
von: Singh, Ajay, et al.
Veröffentlicht: (2026)
von: Singh, Ajay, et al.
Veröffentlicht: (2026)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
von: Gonzalez-Escribano, Arturo, et al.
Veröffentlicht: (2024)
von: Gonzalez-Escribano, Arturo, et al.
Veröffentlicht: (2024)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
Scheduler-Driven Job Atomization
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
A C++17 Thread Pool for High-Performance Scientific Computing
von: Shoshany, Barak
Veröffentlicht: (2021)
von: Shoshany, Barak
Veröffentlicht: (2021)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
SpaDA: A Spatial Dataflow Architecture Programming Language
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
Construction of a Byzantine Linearizable SWMR Atomic Register from SWSR Atomic Registers
von: Kshemkalyani, Ajay D., et al.
Veröffentlicht: (2024)
von: Kshemkalyani, Ajay D., et al.
Veröffentlicht: (2024)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
von: Li, Shigang, et al.
Veröffentlicht: (2020)
von: Li, Shigang, et al.
Veröffentlicht: (2020)
ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
von: Zhu, Xiongwei, et al.
Veröffentlicht: (2026)
von: Zhu, Xiongwei, et al.
Veröffentlicht: (2026)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
von: Lopes, André, et al.
Veröffentlicht: (2024)
von: Lopes, André, et al.
Veröffentlicht: (2024)
NotebookOS: A Replicated Notebook Platform for Interactive Training with On-Demand GPUs
von: Carver, Benjamin, et al.
Veröffentlicht: (2025)
von: Carver, Benjamin, et al.
Veröffentlicht: (2025)
On Advanced Monte Carlo Methods for Linear Algebra on Advanced Accelerator Architectures
von: Lebedev, Anton, et al.
Veröffentlicht: (2024)
von: Lebedev, Anton, et al.
Veröffentlicht: (2024)
Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
Rhizomes and Diffusions for Processing Highly Skewed Graphs on Fine-Grain Message-Driven Systems
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
von: Tramm, John, et al.
Veröffentlicht: (2024) -
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2026) -
Dynamic Memory Management on GPUs with SYCL
von: Standish, Russell K.
Veröffentlicht: (2025) -
Improving inference time in multi-TPU systems with profiled model segmentation
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2025) -
Balanced segmentation of CNNs for multi-TPU inference
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2025)