Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Charles, Bhatia, Sahil, Cheung, Alvin, Shao, Yakun Sophia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Aided Compilation for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2024)
by: Hong, Charles, et al.
Published: (2024)
hdl2v: A Code Translation Dataset for Enhanced LLM Verilog Generation
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
GPU Performance Portability needs Autotuning
by: Ringlein, Burkhard, et al.
Published: (2025)
by: Ringlein, Burkhard, et al.
Published: (2025)
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
by: Lin, Yujun, et al.
Published: (2025)
by: Lin, Yujun, et al.
Published: (2025)
DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
by: Ma, Shaobo, et al.
Published: (2025)
by: Ma, Shaobo, et al.
Published: (2025)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024)
by: Ma, Shaobo, et al.
Published: (2024)
PyraNet: A Multi-Layered Hierarchical Dataset for Verilog
by: Nadimi, Bardia, et al.
Published: (2024)
by: Nadimi, Bardia, et al.
Published: (2024)
Mitigating hallucinations and omissions in LLMs for invertible problems: An application to hardware logic design automation
by: Cassidy, Andrew S., et al.
Published: (2025)
by: Cassidy, Andrew S., et al.
Published: (2025)
VeriMind: Agentic LLM for Automated Verilog Generation with a Novel Evaluation Metric
by: Nadimi, Bardia, et al.
Published: (2025)
by: Nadimi, Bardia, et al.
Published: (2025)
Enabling New HDLs with Agents
by: Zakharov, Mark, et al.
Published: (2024)
by: Zakharov, Mark, et al.
Published: (2024)
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
Zero-Shot RTL Code Generation with Attention Sink Augmented Large Language Models
by: Sandal, Selim, et al.
Published: (2024)
by: Sandal, Selim, et al.
Published: (2024)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
by: Sohn, Gina, et al.
Published: (2025)
by: Sohn, Gina, et al.
Published: (2025)
Evaluating LLMs for Hardware Design and Test
by: Blocklove, Jason, et al.
Published: (2024)
by: Blocklove, Jason, et al.
Published: (2024)
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
Allo: A Programming Model for Composable Accelerator Design
by: Chen, Hongzheng, et al.
Published: (2024)
by: Chen, Hongzheng, et al.
Published: (2024)
Dato: A Task-Based Programming Model for Dataflow Accelerators
by: Fang, Shihan, et al.
Published: (2025)
by: Fang, Shihan, et al.
Published: (2025)
VeriGRAG: Enhancing LLM-Based Verilog Code Generation with Structure-Aware Soft Prompts
by: Zhao, Jiayu, et al.
Published: (2025)
by: Zhao, Jiayu, et al.
Published: (2025)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
by: Meng, Chang, et al.
Published: (2026)
by: Meng, Chang, et al.
Published: (2026)
DecoRTL: A Run-time Decoding Framework for RTL Code Generation with LLMs
by: Akyash, Mohammad, et al.
Published: (2025)
by: Akyash, Mohammad, et al.
Published: (2025)
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
by: Jiang, Wenqi, et al.
Published: (2023)
by: Jiang, Wenqi, et al.
Published: (2023)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
by: Chen, Hongzheng, et al.
Published: (2023)
by: Chen, Hongzheng, et al.
Published: (2023)
RTL++: Graph-enhanced LLM for RTL Code Generation
by: Akyash, Mohammad, et al.
Published: (2025)
by: Akyash, Mohammad, et al.
Published: (2025)
COGNATE: Acceleration of Sparse Tensor Programs on Emerging Hardware using Transfer Learning
by: Sudusinghe, Chamika, et al.
Published: (2025)
by: Sudusinghe, Chamika, et al.
Published: (2025)
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Lorecast: Layout-Aware Performance and Power Forecasting from Natural Language
by: Wang, Runzhi, et al.
Published: (2025)
by: Wang, Runzhi, et al.
Published: (2025)
CoopetitiveV: Leveraging LLM-powered Coopetitive Multi-Agent Prompting for High-quality Verilog Generation
by: Mi, Zhendong, et al.
Published: (2024)
by: Mi, Zhendong, et al.
Published: (2024)
ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning
by: Chen, Zhirong, et al.
Published: (2025)
by: Chen, Zhirong, et al.
Published: (2025)
QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
by: Zhu, Yaoyu, et al.
Published: (2025)
by: Zhu, Yaoyu, et al.
Published: (2025)
'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators
by: Han, Ruichi, et al.
Published: (2026)
by: Han, Ruichi, et al.
Published: (2026)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
by: Liang, Yanbiao, et al.
Published: (2025)
by: Liang, Yanbiao, et al.
Published: (2025)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
by: Wu, Yutong, et al.
Published: (2026)
by: Wu, Yutong, et al.
Published: (2026)
GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and Optimization
by: Abi-Karam, Stefan, et al.
Published: (2023)
by: Abi-Karam, Stefan, et al.
Published: (2023)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
by: Gope, Dibakar, et al.
Published: (2024)
by: Gope, Dibakar, et al.
Published: (2024)
Dynamic Co-Optimization Compiler: Leveraging Multi-Agent Reinforcement Learning for Enhanced DNN Accelerator Performance
by: Fayyazi, Arya, et al.
Published: (2024)
by: Fayyazi, Arya, et al.
Published: (2024)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
Similar Items
-
LLM-Aided Compilation for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2024) -
hdl2v: A Code Translation Dataset for Enhanced LLM Verilog Generation
by: Hong, Charles, et al.
Published: (2025) -
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025) -
GPU Performance Portability needs Autotuning
by: Ringlein, Burkhard, et al.
Published: (2025) -
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
by: Lin, Yujun, et al.
Published: (2025)