T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pati, Suchita, Aga, Shaizeen, Islam, Mahzabeen, Jayasena, Nuwan, Sinclair, Matthew D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Global Optimizations & Lightweight Dynamic Logic for Concurrency
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
von: Pal, Shagnik, et al.
Veröffentlicht: (2025)
von: Pal, Shagnik, et al.
Veröffentlicht: (2025)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
GPU-Augmented OLAP Execution Engine: GPU Offloading
von: Chang, Ilsun
Veröffentlicht: (2025)
von: Chang, Ilsun
Veröffentlicht: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
Moonshot: Optimizing Chain-Based Rotating Leader BFT via Optimistic Proposals
von: Doidge, Isaac, et al.
Veröffentlicht: (2024)
von: Doidge, Isaac, et al.
Veröffentlicht: (2024)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
von: Jung, Myoungsoo
Veröffentlicht: (2025)
von: Jung, Myoungsoo
Veröffentlicht: (2025)
A Survey on Heterogeneous Computing Using SmartNICs and Emerging Data Processing Units
von: Tibbetts, Nathan, et al.
Veröffentlicht: (2025)
von: Tibbetts, Nathan, et al.
Veröffentlicht: (2025)
Konnektor: Connection Protocol for Ensuring Peer Uniqueness in Decentralized P2P Networks
von: Ozkan, Onur
Veröffentlicht: (2024)
von: Ozkan, Onur
Veröffentlicht: (2024)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
von: Dettinger, Falk, et al.
Veröffentlicht: (2025)
von: Dettinger, Falk, et al.
Veröffentlicht: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
Towards Policy-Enabled Multi-Hop Routing for Cross-Chain Message Delivery
von: Rezaei, Amin, et al.
Veröffentlicht: (2026)
von: Rezaei, Amin, et al.
Veröffentlicht: (2026)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
von: Sunkara, Krishna Chaitanya
Veröffentlicht: (2026)
von: Sunkara, Krishna Chaitanya
Veröffentlicht: (2026)
Collective Communication for 100k+ GPUs
von: Si, Min, et al.
Veröffentlicht: (2025)
von: Si, Min, et al.
Veröffentlicht: (2025)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Lincoln AI Computing Survey (LAICS) and Trends
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
Nezha: Deployable and High-Performance Consensus Using Synchronized Clocks
von: Geng, Jinkun, et al.
Veröffentlicht: (2022)
von: Geng, Jinkun, et al.
Veröffentlicht: (2022)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
Security Analysis of Bitcoin's V2 Transport Protocol: Exploiting Design Implications for Sustained Eclipse and Downgrade Attacks
von: Ndolo, Charmaine, et al.
Veröffentlicht: (2026)
von: Ndolo, Charmaine, et al.
Veröffentlicht: (2026)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
von: Bera, Rahul, et al.
Veröffentlicht: (2026)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Global Optimizations & Lightweight Dynamic Logic for Concurrency
von: Pati, Suchita, et al.
Veröffentlicht: (2024) -
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024) -
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
von: Pal, Shagnik, et al.
Veröffentlicht: (2025) -
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025) -
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)