Efficient Orchestrated AI Workflows Execution on Scale-out Spatial Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Jinyi, Tang, Xinru, Yue, Zhiheng, Lu, Guangyang, Yang, Qize, Zhang, Jiahao, Li, Jinxi, Li, Chao, Wei, Shaojun, Hu, Yang, Yin, Shouyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
by: Wang, Huizheng, et al.
Published: (2024)
by: Wang, Huizheng, et al.
Published: (2024)
Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
A Data-Driven Dynamic Execution Orchestration Architecture
by: Bai, Zhenyu, et al.
Published: (2026)
by: Bai, Zhenyu, et al.
Published: (2026)
From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
by: Chen, Xingzhen, et al.
Published: (2026)
by: Chen, Xingzhen, et al.
Published: (2026)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
DAG-aware Synthesis Orchestration
by: Li, Yingjie, et al.
Published: (2023)
by: Li, Yingjie, et al.
Published: (2023)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
by: Jiang, Aojie, et al.
Published: (2026)
by: Jiang, Aojie, et al.
Published: (2026)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
by: Sarda, Giuseppe M., et al.
Published: (2024)
by: Sarda, Giuseppe M., et al.
Published: (2024)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
by: Chen, Paul, et al.
Published: (2026)
by: Chen, Paul, et al.
Published: (2026)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
by: Luo, Shuqing, et al.
Published: (2026)
by: Luo, Shuqing, et al.
Published: (2026)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
by: Li, Tenglong, et al.
Published: (2024)
by: Li, Tenglong, et al.
Published: (2024)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
by: Li, Fang
Published: (2025)
by: Li, Fang
Published: (2025)
NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention Dataflow
by: Chen, Yuehai, et al.
Published: (2025)
by: Chen, Yuehai, et al.
Published: (2025)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM Architecture
by: Chen, Sitian, et al.
Published: (2024)
by: Chen, Sitian, et al.
Published: (2024)
Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
by: Zhou, Weiyu, et al.
Published: (2025)
by: Zhou, Weiyu, et al.
Published: (2025)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
GOMA: Geometrically Optimal Mapping via Analytical Modeling for Spatial Accelerators
by: Yang, Wulve, et al.
Published: (2026)
by: Yang, Wulve, et al.
Published: (2026)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
ERASER: Efficient RTL FAult Simulation Framework with Trimmed Execution Redundancy
by: Tang, Jiaping, et al.
Published: (2025)
by: Tang, Jiaping, et al.
Published: (2025)
Architectural Classification of XR Workloads: Cross-Layer Archetypes and Implications
by: Shi, Xinyu, et al.
Published: (2026)
by: Shi, Xinyu, et al.
Published: (2026)
Reconfigurable Digital RRAM Logic Enables In-Situ Pruning and Learning for Edge AI
by: Wang, Songqi, et al.
Published: (2025)
by: Wang, Songqi, et al.
Published: (2025)
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026)
by: Qian, Chao, et al.
Published: (2026)
LASANA: Large-Scale Surrogate Modeling for Analog Neuromorphic Architecture Exploration
by: Ho, Jason, et al.
Published: (2025)
by: Ho, Jason, et al.
Published: (2025)
Full System Architecture Modeling for Wearable Egocentric Contextual AI
by: Lee, Vincent T., et al.
Published: (2025)
by: Lee, Vincent T., et al.
Published: (2025)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
by: Ge, Mengke, et al.
Published: (2024)
by: Ge, Mengke, et al.
Published: (2024)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
by: Sestito, Cristian, et al.
Published: (2025)
by: Sestito, Cristian, et al.
Published: (2025)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
by: Yi, Xiaoling, et al.
Published: (2026)
by: Yi, Xiaoling, et al.
Published: (2026)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
by: Ding, Hong, et al.
Published: (2025)
by: Ding, Hong, et al.
Published: (2025)
Iterative LLM-Based Assertion Generation Using Syntax-Semantic Representations for Functional Coverage-Guided Verification
by: Wang, Yonghao, et al.
Published: (2026)
by: Wang, Yonghao, et al.
Published: (2026)
LaZagna: An Open-Source Framework for Flexible 3D FPGA Architectural Exploration
by: Youssef, Ismael, et al.
Published: (2025)
by: Youssef, Ismael, et al.
Published: (2025)
Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory
by: Chen, Sitian, et al.
Published: (2026)
by: Chen, Sitian, et al.
Published: (2026)
Similar Items
-
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
by: Wang, Huizheng, et al.
Published: (2024) -
Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
by: Wang, Huizheng, et al.
Published: (2025) -
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
by: Wang, Huizheng, et al.
Published: (2025) -
PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
by: Wang, Huizheng, et al.
Published: (2025) -
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
by: Wang, Huizheng, et al.
Published: (2025)