Global Optimizations & Lightweight Dynamic Logic for Concurrency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pati, Suchita, Aga, Shaizeen, Jayasena, Nuwan, Sinclair, Matthew D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
von: Pal, Shagnik, et al.
Veröffentlicht: (2025)
von: Pal, Shagnik, et al.
Veröffentlicht: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
von: Jeong, Shinnung, et al.
Veröffentlicht: (2025)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
von: Wilhelm, Martin, et al.
Veröffentlicht: (2025)
von: Wilhelm, Martin, et al.
Veröffentlicht: (2025)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
Functionally-Complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
Dynamic Simultaneous Multithreaded Architecture
von: Ortiz-Arroyo, Daniel, et al.
Veröffentlicht: (2024)
von: Ortiz-Arroyo, Daniel, et al.
Veröffentlicht: (2024)
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
von: Green, Conor, et al.
Veröffentlicht: (2024)
von: Green, Conor, et al.
Veröffentlicht: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
von: Bai, Zhenyu, et al.
Veröffentlicht: (2025)
von: Bai, Zhenyu, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
von: Meyer, Marius, et al.
Veröffentlicht: (2024)
von: Meyer, Marius, et al.
Veröffentlicht: (2024)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
Revisiting Computational Storage for Data Integrity and Security
von: Shi, Chao, et al.
Veröffentlicht: (2025)
von: Shi, Chao, et al.
Veröffentlicht: (2025)
Lincoln AI Computing Survey (LAICS) and Trends
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2025)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2025)
MANOJAVAM: A Scalable, Unified FPGA Accelerator for Matrix Multiplication and Singular Value Decomposition in Principal Component Analysis
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2026)
von: Villarrubia, Jorge, et al.
Veröffentlicht: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
von: Pati, Suchita, et al.
Veröffentlicht: (2024) -
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024) -
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025) -
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025) -
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
von: Pal, Shagnik, et al.
Veröffentlicht: (2025)