A System Level Compiler for Massively-Parallel, Spatial, Dataflow Architectures
Fuente:
arXiv
Salvato in:
| Autori principali: | Van Essendelft, Dirk, Wingo, Patrick, Jordan, Terry, Smith, Ryan, Saidi, Wissam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
di: Zhang, Xu, et al.
Pubblicazione: (2024)
di: Zhang, Xu, et al.
Pubblicazione: (2024)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
di: Call, Aaron, et al.
Pubblicazione: (2025)
di: Call, Aaron, et al.
Pubblicazione: (2025)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
di: Vo, Huynh Q. N., et al.
Pubblicazione: (2025)
di: Vo, Huynh Q. N., et al.
Pubblicazione: (2025)
CLAASIC: a Cortex-Inspired Hardware Accelerator
di: Puente, Valentin, et al.
Pubblicazione: (2016)
di: Puente, Valentin, et al.
Pubblicazione: (2016)
Reference Architecture of a Quantum-Centric Supercomputer
di: Seelam, Seetharami, et al.
Pubblicazione: (2026)
di: Seelam, Seetharami, et al.
Pubblicazione: (2026)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
di: Tsai, Yu-Shiang, et al.
Pubblicazione: (2024)
di: Tsai, Yu-Shiang, et al.
Pubblicazione: (2024)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
di: Ortega, Cristobal, et al.
Pubblicazione: (2024)
di: Ortega, Cristobal, et al.
Pubblicazione: (2024)
Kitsune: Enabling Dataflow Execution on GPUs
di: Davies, Michael, et al.
Pubblicazione: (2025)
di: Davies, Michael, et al.
Pubblicazione: (2025)
INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation Processing
di: Abi-Karam, Stefan, et al.
Pubblicazione: (2023)
di: Abi-Karam, Stefan, et al.
Pubblicazione: (2023)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
di: Shi, Man, et al.
Pubblicazione: (2024)
di: Shi, Man, et al.
Pubblicazione: (2024)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2026)
di: Afzal, Ayesha, et al.
Pubblicazione: (2026)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
di: Negi, Shubham, et al.
Pubblicazione: (2025)
di: Negi, Shubham, et al.
Pubblicazione: (2025)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
di: Szafarczyk, Robert, et al.
Pubblicazione: (2025)
di: Szafarczyk, Robert, et al.
Pubblicazione: (2025)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
di: Park, Jihoon, et al.
Pubblicazione: (2025)
di: Park, Jihoon, et al.
Pubblicazione: (2025)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
di: Hou, Yuewen, et al.
Pubblicazione: (2025)
di: Hou, Yuewen, et al.
Pubblicazione: (2025)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
di: Filippov, Dmitry, et al.
Pubblicazione: (2025)
di: Filippov, Dmitry, et al.
Pubblicazione: (2025)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
di: Roy, Rohan Basu, et al.
Pubblicazione: (2025)
di: Roy, Rohan Basu, et al.
Pubblicazione: (2025)
Carbon Connect: An Ecosystem for Sustainable Computing
di: Lee, Benjamin C., et al.
Pubblicazione: (2024)
di: Lee, Benjamin C., et al.
Pubblicazione: (2024)
Managed-Retention Memory: A New Class of Memory for the AI Era
di: Legtchenko, Sergey, et al.
Pubblicazione: (2025)
di: Legtchenko, Sergey, et al.
Pubblicazione: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
Parendi: Thousand-Way Parallel RTL Simulation
di: Emami, Mahyar, et al.
Pubblicazione: (2024)
di: Emami, Mahyar, et al.
Pubblicazione: (2024)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
di: Vo, Huynh Q. N., et al.
Pubblicazione: (2025)
di: Vo, Huynh Q. N., et al.
Pubblicazione: (2025)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
di: Tithi, Jesmin Jahan, et al.
Pubblicazione: (2025)
di: Tithi, Jesmin Jahan, et al.
Pubblicazione: (2025)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
di: Garg, Chirag, et al.
Pubblicazione: (2025)
di: Garg, Chirag, et al.
Pubblicazione: (2025)
General-Purpose Multicore Architectures
di: Ghose, Saugata
Pubblicazione: (2024)
di: Ghose, Saugata
Pubblicazione: (2024)
Dynamic Simultaneous Multithreaded Architecture
di: Ortiz-Arroyo, Daniel, et al.
Pubblicazione: (2024)
di: Ortiz-Arroyo, Daniel, et al.
Pubblicazione: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
di: Luo, Weile, et al.
Pubblicazione: (2025)
di: Luo, Weile, et al.
Pubblicazione: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
di: Li, Jonathan, et al.
Pubblicazione: (2025)
di: Li, Jonathan, et al.
Pubblicazione: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
PIUMA: Programmable Integrated Unified Memory Architecture
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
Efficient Architecture for RISC-V Vector Memory Access
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
di: Zhu, Yu, et al.
Pubblicazione: (2025)
di: Zhu, Yu, et al.
Pubblicazione: (2025)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
di: Pan, Xueting, et al.
Pubblicazione: (2024)
di: Pan, Xueting, et al.
Pubblicazione: (2024)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
Documenti analoghi
-
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
di: Zhang, Xu, et al.
Pubblicazione: (2024) -
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
di: Call, Aaron, et al.
Pubblicazione: (2025) -
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
di: Vo, Huynh Q. N., et al.
Pubblicazione: (2025) -
CLAASIC: a Cortex-Inspired Hardware Accelerator
di: Puente, Valentin, et al.
Pubblicazione: (2016) -
Reference Architecture of a Quantum-Centric Supercomputer
di: Seelam, Seetharami, et al.
Pubblicazione: (2026)