COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Negi, Shubham, Singhal, Manik, Ankit, Aayush, Bhoja, Sudeep, Roy, Kaushik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
COMET: Neural Cost Model Explanation Framework
von: Chaudhary, Isha, et al.
Veröffentlicht: (2023)
von: Chaudhary, Isha, et al.
Veröffentlicht: (2023)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
Cloud-Native Operation of Roadside Infrastructure Enabling Demand-Driven Collective Perception via V2X
von: Zanger, Lukas, et al.
Veröffentlicht: (2026)
von: Zanger, Lukas, et al.
Veröffentlicht: (2026)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
von: Noh, Si Ung, et al.
Veröffentlicht: (2024)
von: Noh, Si Ung, et al.
Veröffentlicht: (2024)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
von: Garg, Raveesh, et al.
Veröffentlicht: (2023)
von: Garg, Raveesh, et al.
Veröffentlicht: (2023)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
EPOCH: Enabling Preemption Operation for Context Saving in Heterogeneous FPGA Systems
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
WWW: What, When, Where to Compute-in-Memory
von: Sharma, Tanvi, et al.
Veröffentlicht: (2023)
von: Sharma, Tanvi, et al.
Veröffentlicht: (2023)
Muchisim: A Simulation Framework for Design Exploration of Multi-Chip Manycore Systems
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation Processing
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2023)
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2023)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
von: Green, Conor, et al.
Veröffentlicht: (2024)
von: Green, Conor, et al.
Veröffentlicht: (2024)
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
von: Herbst, Steven, et al.
Veröffentlicht: (2024)
von: Herbst, Steven, et al.
Veröffentlicht: (2024)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Pooling Engram Conditional Memory in Large Language Models using CXL
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
von: Oliveira, Geraldo F.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F.
Veröffentlicht: (2025)
Enhancing Regression Models for Complex Systems Using Evolutionary Techniques for Feature Engineering
von: Arroba, Patricia, et al.
Veröffentlicht: (2024)
von: Arroba, Patricia, et al.
Veröffentlicht: (2024)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
von: Tian, Yuyang, et al.
Veröffentlicht: (2025)
von: Tian, Yuyang, et al.
Veröffentlicht: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
A Modern Primer on Processing in Memory
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
The Dawn of Disaggregation and the Coherence Conundrum: A Call for Federated Coherence
von: Hong, Jaewan, et al.
Veröffentlicht: (2025)
von: Hong, Jaewan, et al.
Veröffentlicht: (2025)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025) -
COMET: Neural Cost Model Explanation Framework
von: Chaudhary, Isha, et al.
Veröffentlicht: (2023) -
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024) -
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023) -
Cloud-Native Operation of Roadside Infrastructure Enabling Demand-Driven Collective Perception via V2X
von: Zanger, Lukas, et al.
Veröffentlicht: (2026)