Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | De Sensi, Daniele, Pichetti, Lorenzo, Vella, Flavio, De Matteis, Tiziano, Ren, Zebin, Fusco, Luigi, Turisini, Matteo, Cesarini, Daniele, Lust, Kurt, Trivedi, Animesh, Roweth, Duncan, Spiga, Filippo, Di Girolamo, Salvatore, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Injection-Throttling Mechanisms for Congestion Control for Data-center and Supercomputer Interconnects
von: Olmedilla, Cristina, et al.
Veröffentlicht: (2025)
von: Olmedilla, Cristina, et al.
Veröffentlicht: (2025)
On the Power Saving in High-Speed Ethernet-based Networks for Supercomputers and Data Centers
von: de la Rosa, Miguel Sánchez, et al.
Veröffentlicht: (2025)
von: de la Rosa, Miguel Sánchez, et al.
Veröffentlicht: (2025)
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
von: Özyılmaz, Mustafa Mert
Veröffentlicht: (2026)
von: Özyılmaz, Mustafa Mert
Veröffentlicht: (2026)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
von: Rout, Nikhil, et al.
Veröffentlicht: (2025)
von: Rout, Nikhil, et al.
Veröffentlicht: (2025)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
von: Dudek, Piotr
Veröffentlicht: (2024)
von: Dudek, Piotr
Veröffentlicht: (2024)
On the Benefits of Traffic "Reprofiling" -- The Multiple Hops Case -- Part I
von: Qiu, Jiaming, et al.
Veröffentlicht: (2024)
von: Qiu, Jiaming, et al.
Veröffentlicht: (2024)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
von: Kang, Hyunwoo
Veröffentlicht: (2026)
von: Kang, Hyunwoo
Veröffentlicht: (2026)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
Lincoln AI Computing Survey (LAICS) and Trends
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
Orchestrating Data Collection and Computation in Green IoT Networks
von: Zhan, Junfei, et al.
Veröffentlicht: (2026)
von: Zhan, Junfei, et al.
Veröffentlicht: (2026)
Cost-effective and performant virtual WANs with CORNIFER
von: Anjali, et al.
Veröffentlicht: (2024)
von: Anjali, et al.
Veröffentlicht: (2024)
BECS: A Privacy-Preserving Computing Resource Sharing Mechanism for 6G Computing Power Network
von: Yan, Kun, et al.
Veröffentlicht: (2024)
von: Yan, Kun, et al.
Veröffentlicht: (2024)
On the Benefits of Traffic "Reprofiling" -- The Multiple Hops Case -- Part II
von: Qiu, Jiaming, et al.
Veröffentlicht: (2026)
von: Qiu, Jiaming, et al.
Veröffentlicht: (2026)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
von: Giorgi, Roberto, et al.
Veröffentlicht: (2026)
von: Giorgi, Roberto, et al.
Veröffentlicht: (2026)
GPU-Initiated Networking for NCCL
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
von: Li, Richie
Veröffentlicht: (2025)
von: Li, Richie
Veröffentlicht: (2025)
Solving the Problem of Poor Internet Connectivity in Dhaka: Innovative Solutions Using Advanced WebRTC and Adaptive Streaming Technologies
von: Malinovskiy, Pavel
Veröffentlicht: (2025)
von: Malinovskiy, Pavel
Veröffentlicht: (2025)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
von: De Sensi, Daniele, et al.
Veröffentlicht: (2025)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2025)
Real-Time Aware IP-Networking for Resource-Constrained Embedded Devices
von: Behnke, Ilja
Veröffentlicht: (2024)
von: Behnke, Ilja
Veröffentlicht: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
von: Palaniappan, Kathiravan
Veröffentlicht: (2026)
von: Palaniappan, Kathiravan
Veröffentlicht: (2026)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
von: Liu, Lian, et al.
Veröffentlicht: (2025)
von: Liu, Lian, et al.
Veröffentlicht: (2025)
Nezha: Deployable and High-Performance Consensus Using Synchronized Clocks
von: Geng, Jinkun, et al.
Veröffentlicht: (2022)
von: Geng, Jinkun, et al.
Veröffentlicht: (2022)
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
von: Behrendt, Emet, et al.
Veröffentlicht: (2025)
von: Behrendt, Emet, et al.
Veröffentlicht: (2025)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
von: Jung, Myoungsoo
Veröffentlicht: (2025)
von: Jung, Myoungsoo
Veröffentlicht: (2025)
CEO-DC: Driving Decarbonization in HPC Data Centers with Actionable Insights
von: Álvarez, Rubén Rodríguez, et al.
Veröffentlicht: (2025)
von: Álvarez, Rubén Rodríguez, et al.
Veröffentlicht: (2025)
Traffic-Aware Configuration of OPC UA PubSub in Industrial Automation Networks
von: Ekrad, Kasra, et al.
Veröffentlicht: (2026)
von: Ekrad, Kasra, et al.
Veröffentlicht: (2026)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
Characterizing 5G User Throughput via Uncertainty Modeling and Crowdsourced Measurements
von: Albert-Smet, Javier, et al.
Veröffentlicht: (2025)
von: Albert-Smet, Javier, et al.
Veröffentlicht: (2025)
LLM-Driven Large-Scale Spectrum Access
von: Yang, Ning, et al.
Veröffentlicht: (2026)
von: Yang, Ning, et al.
Veröffentlicht: (2026)
Enabling full-speed random access to the entire memory on the A100 GPU
von: Walker, Alden
Veröffentlicht: (2024)
von: Walker, Alden
Veröffentlicht: (2024)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
Fast and Fusiest: An Optimal Fusion-Aware Mapper for Accelerator Design
von: Andrulis, Tanner, et al.
Veröffentlicht: (2026)
von: Andrulis, Tanner, et al.
Veröffentlicht: (2026)
The Turbo-Charged Mapper: Fast and Optimal Mapping for Energy-efficient and Low-latency Accelerator Design
von: Gilbert, Michael, et al.
Veröffentlicht: (2026)
von: Gilbert, Michael, et al.
Veröffentlicht: (2026)
3GPP NR V2X Mode 2d: Analysis of Distributed Scheduling for Groupcast using ns-3 5G LENA Simulator
von: Fehrenbach, Thomas, et al.
Veröffentlicht: (2025)
von: Fehrenbach, Thomas, et al.
Veröffentlicht: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
A Per-Access Upper Bound for Shared-Resource Interference in Direct-Mapped Multicore Architectures
von: Pedroni, Felipe T.
Veröffentlicht: (2026)
von: Pedroni, Felipe T.
Veröffentlicht: (2026)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
von: Mo, Zhiwen, et al.
Veröffentlicht: (2024)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving Injection-Throttling Mechanisms for Congestion Control for Data-center and Supercomputer Interconnects
von: Olmedilla, Cristina, et al.
Veröffentlicht: (2025) -
On the Power Saving in High-Speed Ethernet-based Networks for Supercomputers and Data Centers
von: de la Rosa, Miguel Sánchez, et al.
Veröffentlicht: (2025) -
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024) -
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
von: Özyılmaz, Mustafa Mert
Veröffentlicht: (2026) -
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)