ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Zhongkai, Zhou, Chenyang, Lin, Yichen, Zhang, Hejia, Ye, Haotian, Cui, Junxia, Pan, Zaifeng, Zhao, Jishen, Ding, Yufei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915766320758784
author Yu, Zhongkai
Zhou, Chenyang
Lin, Yichen
Zhang, Hejia
Ye, Haotian
Cui, Junxia
Pan, Zaifeng
Zhao, Jishen
Ding, Yufei
author_facet Yu, Zhongkai
Zhou, Chenyang
Lin, Yichen
Zhang, Hejia
Ye, Haotian
Cui, Junxia
Pan, Zaifeng
Zhao, Jishen
Ding, Yufei
contents While Large Language Models (LLMs) show significant potential in hardware engineering, current benchmarks suffer from saturation and limited task diversity, failing to reflect LLMs' performance in real industrial workflows. To address this gap, we propose a comprehensive benchmark for AI-aided chip design that rigorously evaluates LLMs across three critical tasks: Verilog generation, debugging, and reference model generation. Our benchmark features 44 realistic modules with complex hierarchical structures, 89 systematic debugging cases, and 132 reference model samples across Python, SystemC, and CXXRTL. Evaluation results reveal substantial performance gaps, with state-of-the-art Claude-4.5-opus achieving only 30.74\% on Verilog generation and 13.33\% on Python reference model generation, demonstrating significant challenges compared to existing saturated benchmarks where SOTA models achieve over 95\% pass rates. Additionally, to help enhance LLM reference model generation, we provide an automated toolbox for high-quality training data generation, facilitating future research in this underexplored domain. Our code is available at https://github.com/zhongkaiyu/ChipBench.git.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21448
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
Yu, Zhongkai
Zhou, Chenyang
Lin, Yichen
Zhang, Hejia
Ye, Haotian
Cui, Junxia
Pan, Zaifeng
Zhao, Jishen
Ding, Yufei
Artificial Intelligence
Hardware Architecture
While Large Language Models (LLMs) show significant potential in hardware engineering, current benchmarks suffer from saturation and limited task diversity, failing to reflect LLMs' performance in real industrial workflows. To address this gap, we propose a comprehensive benchmark for AI-aided chip design that rigorously evaluates LLMs across three critical tasks: Verilog generation, debugging, and reference model generation. Our benchmark features 44 realistic modules with complex hierarchical structures, 89 systematic debugging cases, and 132 reference model samples across Python, SystemC, and CXXRTL. Evaluation results reveal substantial performance gaps, with state-of-the-art Claude-4.5-opus achieving only 30.74\% on Verilog generation and 13.33\% on Python reference model generation, demonstrating significant challenges compared to existing saturated benchmarks where SOTA models achieve over 95\% pass rates. Additionally, to help enhance LLM reference model generation, we provide an automated toolbox for high-quality training data generation, facilitating future research in this underexplored domain. Our code is available at https://github.com/zhongkaiyu/ChipBench.git.
title ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
topic Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2601.21448