LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Yujun, Zhang, Zhekai, Han, Song |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
di: Ma, Shaobo, et al.
Pubblicazione: (2025)
di: Ma, Shaobo, et al.
Pubblicazione: (2025)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
di: Ma, Shaobo, et al.
Pubblicazione: (2024)
di: Ma, Shaobo, et al.
Pubblicazione: (2024)
GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and Optimization
di: Abi-Karam, Stefan, et al.
Pubblicazione: (2023)
di: Abi-Karam, Stefan, et al.
Pubblicazione: (2023)
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
di: Hong, Charles, et al.
Pubblicazione: (2025)
di: Hong, Charles, et al.
Pubblicazione: (2025)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
di: Lu, Jinming, et al.
Pubblicazione: (2025)
di: Lu, Jinming, et al.
Pubblicazione: (2025)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
di: Cong, Rongqing, et al.
Pubblicazione: (2024)
di: Cong, Rongqing, et al.
Pubblicazione: (2024)
Dynamic Co-Optimization Compiler: Leveraging Multi-Agent Reinforcement Learning for Enhanced DNN Accelerator Performance
di: Fayyazi, Arya, et al.
Pubblicazione: (2024)
di: Fayyazi, Arya, et al.
Pubblicazione: (2024)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators
di: Han, Ruichi, et al.
Pubblicazione: (2026)
di: Han, Ruichi, et al.
Pubblicazione: (2026)
COGNATE: Acceleration of Sparse Tensor Programs on Emerging Hardware using Transfer Learning
di: Sudusinghe, Chamika, et al.
Pubblicazione: (2025)
di: Sudusinghe, Chamika, et al.
Pubblicazione: (2025)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
di: Chen, Hongzheng, et al.
Pubblicazione: (2023)
di: Chen, Hongzheng, et al.
Pubblicazione: (2023)
OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection
di: Cui, Fan, et al.
Pubblicazione: (2024)
di: Cui, Fan, et al.
Pubblicazione: (2024)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
di: Zhang, Jiahao, et al.
Pubblicazione: (2026)
di: Zhang, Jiahao, et al.
Pubblicazione: (2026)
MATADOR: Automated System-on-Chip Tsetlin Machine Design Generation for Edge Applications
di: Rahman, Tousif, et al.
Pubblicazione: (2024)
di: Rahman, Tousif, et al.
Pubblicazione: (2024)
LLM-Enhanced Bayesian Optimization for Efficient Analog Layout Constraint Generation
di: Chen, Guojin, et al.
Pubblicazione: (2024)
di: Chen, Guojin, et al.
Pubblicazione: (2024)
Heterogeneous Acceleration Pipeline for Recommendation System Training
di: Adnan, Muhammad, et al.
Pubblicazione: (2022)
di: Adnan, Muhammad, et al.
Pubblicazione: (2022)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
di: Ding, Ruogu, et al.
Pubblicazione: (2025)
di: Ding, Ruogu, et al.
Pubblicazione: (2025)
MG-Verilog: Multi-grained Dataset Towards Enhanced LLM-assisted Verilog Generation
di: Zhang, Yongan, et al.
Pubblicazione: (2024)
di: Zhang, Yongan, et al.
Pubblicazione: (2024)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
di: Demirkiran, Cansu, et al.
Pubblicazione: (2023)
di: Demirkiran, Cansu, et al.
Pubblicazione: (2023)
LaMAGIC2: Advanced Circuit Formulations for Language Model-Based Analog Topology Generation
di: Chang, Chen-Chia, et al.
Pubblicazione: (2025)
di: Chang, Chen-Chia, et al.
Pubblicazione: (2025)
The Role of Advanced Computer Architectures in Accelerating Artificial Intelligence Workloads
di: Amin, Shahid, et al.
Pubblicazione: (2025)
di: Amin, Shahid, et al.
Pubblicazione: (2025)
Differentiable Initialization-Accelerated CPU-GPU Hybrid Combinatorial Scheduling
di: Liu, Mingju, et al.
Pubblicazione: (2026)
di: Liu, Mingju, et al.
Pubblicazione: (2026)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
di: Xu, Bin, et al.
Pubblicazione: (2025)
di: Xu, Bin, et al.
Pubblicazione: (2025)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
di: Kabir, Ehsan, et al.
Pubblicazione: (2024)
di: Kabir, Ehsan, et al.
Pubblicazione: (2024)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
di: Meng, Chang, et al.
Pubblicazione: (2026)
di: Meng, Chang, et al.
Pubblicazione: (2026)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
di: Qiao, Ye, et al.
Pubblicazione: (2026)
di: Qiao, Ye, et al.
Pubblicazione: (2026)
AutoHLS: Learning to Accelerate Design Space Exploration for HLS Designs
di: Ahmed, Md Rubel, et al.
Pubblicazione: (2024)
di: Ahmed, Md Rubel, et al.
Pubblicazione: (2024)
Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL
di: Cai, Siyang, et al.
Pubblicazione: (2026)
di: Cai, Siyang, et al.
Pubblicazione: (2026)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
AdAM: Adaptive Fault-Tolerant Approximate Multiplier for Edge DNN Accelerators
di: Taheri, Mahdi, et al.
Pubblicazione: (2024)
di: Taheri, Mahdi, et al.
Pubblicazione: (2024)
SAFFIRA: a Framework for Assessing the Reliability of Systolic-Array-Based DNN Accelerators
di: Taheri, Mahdi, et al.
Pubblicazione: (2024)
di: Taheri, Mahdi, et al.
Pubblicazione: (2024)
Agent Factories for High Level Synthesis: How Far Can General-Purpose Coding Agents Go in Hardware Optimization?
di: Bhandwaldar, Abhishek, et al.
Pubblicazione: (2026)
di: Bhandwaldar, Abhishek, et al.
Pubblicazione: (2026)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2025)
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
di: Lübeck, Konstantin, et al.
Pubblicazione: (2024)
di: Lübeck, Konstantin, et al.
Pubblicazione: (2024)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
di: Li, Guoyu, et al.
Pubblicazione: (2025)
di: Li, Guoyu, et al.
Pubblicazione: (2025)
AIRCHITECT v2: Learning the Hardware Accelerator Design Space through Unified Representations
di: Seo, Jamin, et al.
Pubblicazione: (2025)
di: Seo, Jamin, et al.
Pubblicazione: (2025)
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
di: Helal, Shereef, et al.
Pubblicazione: (2025)
di: Helal, Shereef, et al.
Pubblicazione: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
di: Kim, Jiyoon, et al.
Pubblicazione: (2025)
di: Kim, Jiyoon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
di: Wang, Hanrui, et al.
Pubblicazione: (2020) -
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
di: Ma, Shaobo, et al.
Pubblicazione: (2025) -
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
di: Ma, Shaobo, et al.
Pubblicazione: (2024) -
GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and Optimization
di: Abi-Karam, Stefan, et al.
Pubblicazione: (2023) -
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
di: Hong, Charles, et al.
Pubblicazione: (2025)