Allo: A Programming Model for Composable Accelerator Design
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Hongzheng, Zhang, Niansong, Xiang, Shaojie, Zeng, Zhichen, Dai, Mengjia, Zhang, Zhiru |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dato: A Task-Based Programming Model for Dataflow Accelerators
by: Fang, Shihan, et al.
Published: (2025)
by: Fang, Shihan, et al.
Published: (2025)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
by: Chen, Hongzheng, et al.
Published: (2023)
by: Chen, Hongzheng, et al.
Published: (2023)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
by: Sohn, Gina, et al.
Published: (2025)
by: Sohn, Gina, et al.
Published: (2025)
LLM-Aided Compilation for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2024)
by: Hong, Charles, et al.
Published: (2024)
QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpression for Circuit Design
by: Huang, Lei, et al.
Published: (2025)
by: Huang, Lei, et al.
Published: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
by: Zhang, Niansong, et al.
Published: (2025)
by: Zhang, Niansong, et al.
Published: (2025)
Leveraging High-Level Synthesis and Large Language Models to Generate, Simulate, and Deploy a Uniform Random Number Generator Hardware Design
by: Meech, James T.
Published: (2023)
by: Meech, James T.
Published: (2023)
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
by: Zhu, Yaoyu, et al.
Published: (2025)
by: Zhu, Yaoyu, et al.
Published: (2025)
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
by: Lacouture, Rubens, et al.
Published: (2025)
by: Lacouture, Rubens, et al.
Published: (2025)
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
DecoRTL: A Run-time Decoding Framework for RTL Code Generation with LLMs
by: Akyash, Mohammad, et al.
Published: (2025)
by: Akyash, Mohammad, et al.
Published: (2025)
RTL++: Graph-enhanced LLM for RTL Code Generation
by: Akyash, Mohammad, et al.
Published: (2025)
by: Akyash, Mohammad, et al.
Published: (2025)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024)
by: Zhang, Weichuang, et al.
Published: (2024)
GauS: Differentiable Scheduling Optimization via Gaussian Reparameterization
by: Cai, Yaohui, et al.
Published: (2026)
by: Cai, Yaohui, et al.
Published: (2026)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Partial Cross-Compilation and Mixed Execution for Accelerating Dynamic Binary Translation
by: Gu, Yuhao, et al.
Published: (2025)
by: Gu, Yuhao, et al.
Published: (2025)
Evaluating LLMs for Hardware Design and Test
by: Blocklove, Jason, et al.
Published: (2024)
by: Blocklove, Jason, et al.
Published: (2024)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution
by: Liu, Shang, et al.
Published: (2023)
by: Liu, Shang, et al.
Published: (2023)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)
by: Wang, Erwei, et al.
Published: (2025)
PyraNet: A Multi-Layered Hierarchical Dataset for Verilog
by: Nadimi, Bardia, et al.
Published: (2024)
by: Nadimi, Bardia, et al.
Published: (2024)
GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
hdl2v: A Code Translation Dataset for Enhanced LLM Verilog Generation
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
Less is More: Hop-Wise Graph Attention for Scalable and Generalizable Learning on Circuits
by: Deng, Chenhui, et al.
Published: (2024)
by: Deng, Chenhui, et al.
Published: (2024)
Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
Relational Hoare Logic for High-Level Synthesis of Hardware Accelerators
by: Tanaka, Izumi, et al.
Published: (2026)
by: Tanaka, Izumi, et al.
Published: (2026)
Optimizing High-Level Synthesis Designs with Retrieval-Augmented Large Language Models
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration
by: Zhuang, Jinming, et al.
Published: (2024)
by: Zhuang, Jinming, et al.
Published: (2024)
MINISA: Minimal Instruction Set Architecture for Next-gen Reconfigurable Inference Accelerator
by: Tong, Jianming, et al.
Published: (2026)
by: Tong, Jianming, et al.
Published: (2026)
ACT: Automatically Generating Compiler Backends from Tensor Accelerator ISA Descriptions
by: Jain, Devansh, et al.
Published: (2025)
by: Jain, Devansh, et al.
Published: (2025)
Enabling New HDLs with Agents
by: Zakharov, Mark, et al.
Published: (2024)
by: Zakharov, Mark, et al.
Published: (2024)
Mitigating hallucinations and omissions in LLMs for invertible problems: An application to hardware logic design automation
by: Cassidy, Andrew S., et al.
Published: (2025)
by: Cassidy, Andrew S., et al.
Published: (2025)
VeriMind: Agentic LLM for Automated Verilog Generation with a Novel Evaluation Metric
by: Nadimi, Bardia, et al.
Published: (2025)
by: Nadimi, Bardia, et al.
Published: (2025)
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
GraCo -- A Graph Composer for Integrated Circuits
by: Uhlich, Stefan, et al.
Published: (2024)
by: Uhlich, Stefan, et al.
Published: (2024)
Similar Items
-
Dato: A Task-Based Programming Model for Dataflow Accelerators
by: Fang, Shihan, et al.
Published: (2025) -
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
by: Chen, Hongzheng, et al.
Published: (2023) -
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
by: Chen, Hongzheng, et al.
Published: (2025) -
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
by: Sohn, Gina, et al.
Published: (2025) -
LLM-Aided Compilation for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2024)