In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hoefler, Torsten, Khalilov, Mikhail, Clark, Josiah, Anubolu, Surendra, Kalkunte, Mohan, Schramm, Karen, Spada, Eric, Roweth, Duncan, Underwood, Keith, Caulfield, Adrian, Kabbani, Abdul, Rastegari, Amirreza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ultra Ethernet's Design Principles and Architectural Innovations
von: Hoefler, Torsten, et al.
Veröffentlicht: (2025)
von: Hoefler, Torsten, et al.
Veröffentlicht: (2025)
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal
von: Jiang, Wenqi, et al.
Veröffentlicht: (2024)
von: Jiang, Wenqi, et al.
Veröffentlicht: (2024)
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
von: Riedel, Samuel, et al.
Veröffentlicht: (2024)
von: Riedel, Samuel, et al.
Veröffentlicht: (2024)
PlaceIT: Placement-based Inter-Chiplet Interconnect Topologies
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
von: Iff, Patrick, et al.
Veröffentlicht: (2026)
von: Iff, Patrick, et al.
Veröffentlicht: (2026)
Hazel: Secure and Efficient Disaggregated Storage
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
von: Iff, Patrick, et al.
Veröffentlicht: (2023)
von: Iff, Patrick, et al.
Veröffentlicht: (2023)
Workload Characterization for Branch Predictability
von: Vikas, FNU, et al.
Veröffentlicht: (2025)
von: Vikas, FNU, et al.
Veröffentlicht: (2025)
Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
von: Popovici, Doru Thom, et al.
Veröffentlicht: (2025)
von: Popovici, Doru Thom, et al.
Veröffentlicht: (2025)
Architectural Classification of XR Workloads: Cross-Layer Archetypes and Implications
von: Shi, Xinyu, et al.
Veröffentlicht: (2026)
von: Shi, Xinyu, et al.
Veröffentlicht: (2026)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
von: Ge, Mengke, et al.
Veröffentlicht: (2024)
von: Ge, Mengke, et al.
Veröffentlicht: (2024)
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
Communication Characterization of AI Workloads for Large-scale Multi-chiplet Accelerators
von: Musavi, Mariam, et al.
Veröffentlicht: (2024)
von: Musavi, Mariam, et al.
Veröffentlicht: (2024)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
von: Purayil, Navaneeth Kunhi, et al.
Veröffentlicht: (2025)
von: Purayil, Navaneeth Kunhi, et al.
Veröffentlicht: (2025)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
Messaging-based Adaptive Vector Computing (MAVeC) Accelerator for AI Workloads
von: Chowdhury, Md. Rownak Hossain, et al.
Veröffentlicht: (2024)
von: Chowdhury, Md. Rownak Hossain, et al.
Veröffentlicht: (2024)
CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures
von: Qi, Yingjie, et al.
Veröffentlicht: (2025)
von: Qi, Yingjie, et al.
Veröffentlicht: (2025)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
von: Li, Boyu, et al.
Veröffentlicht: (2025)
von: Li, Boyu, et al.
Veröffentlicht: (2025)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
von: Hayes, Oran, et al.
Veröffentlicht: (2026)
von: Hayes, Oran, et al.
Veröffentlicht: (2026)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
von: Sestito, Cristian, et al.
Veröffentlicht: (2025)
von: Sestito, Cristian, et al.
Veröffentlicht: (2025)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
von: Khabbazan, Bahareh, et al.
Veröffentlicht: (2025)
von: Khabbazan, Bahareh, et al.
Veröffentlicht: (2025)
Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-based Accelerators
von: Das, Abhijit, et al.
Veröffentlicht: (2022)
von: Das, Abhijit, et al.
Veröffentlicht: (2022)
HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads
von: Negi, Shubham, et al.
Veröffentlicht: (2024)
von: Negi, Shubham, et al.
Veröffentlicht: (2024)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
DCI: A Coordinated Allocation and Filling Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
von: Luo, Yi, et al.
Veröffentlicht: (2025)
von: Luo, Yi, et al.
Veröffentlicht: (2025)
Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
von: Kwon, Jaewon, et al.
Veröffentlicht: (2025)
von: Kwon, Jaewon, et al.
Veröffentlicht: (2025)
No One-Size-Fits-All: A Workload-Driven Characterization of Bit-Parallel vs. Bit-Serial Data Layouts for Processing-using-Memory
von: Zhang, Jingyao, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyao, et al.
Veröffentlicht: (2025)
WaSP: Warp Scheduling to Mimic Prefetching in Graphics Workloads
von: Joseph, Diya, et al.
Veröffentlicht: (2024)
von: Joseph, Diya, et al.
Veröffentlicht: (2024)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
von: Vahdatniya, Parmida, et al.
Veröffentlicht: (2025)
von: Vahdatniya, Parmida, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
von: Zhang, Niansong, et al.
Veröffentlicht: (2025)
von: Zhang, Niansong, et al.
Veröffentlicht: (2025)
FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
von: Sali, Safa Mohammed, et al.
Veröffentlicht: (2025)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
von: Feng, Yinxiao, et al.
Veröffentlicht: (2025)
von: Feng, Yinxiao, et al.
Veröffentlicht: (2025)
CapsBeam: Accelerating Capsule Network based Beamformer for Ultrasound Non-Steered Plane Wave Imaging on Field Programmable Gate Array
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ultra Ethernet's Design Principles and Architectural Innovations
von: Hoefler, Torsten, et al.
Veröffentlicht: (2025) -
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
von: Iff, Patrick, et al.
Veröffentlicht: (2025) -
Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal
von: Jiang, Wenqi, et al.
Veröffentlicht: (2024) -
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
von: Riedel, Samuel, et al.
Veröffentlicht: (2024) -
PlaceIT: Placement-based Inter-Chiplet Interconnect Topologies
von: Iff, Patrick, et al.
Veröffentlicht: (2025)