RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shang, Fang, Wenji, Lu, Yao, Zhang, Qijun, Zhang, Hongce, Xie, Zhiyao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915430909607936
author Liu, Shang
Fang, Wenji
Lu, Yao
Zhang, Qijun
Zhang, Hongce
Xie, Zhiyao
author_facet Liu, Shang
Fang, Wenji
Lu, Yao
Zhang, Qijun
Zhang, Hongce
Xie, Zhiyao
contents The automatic generation of RTL code (e.g., Verilog) using natural language instructions and large language models (LLMs) has attracted significant research interest recently. However, most existing approaches heavily rely on commercial LLMs such as ChatGPT, while open-source LLMs tailored for this specific design generation task exhibit notably inferior performance. The absence of high-quality open-source solutions restricts the flexibility and data privacy of this emerging technique. In this study, we present a new customized LLM solution with a modest parameter count of only 7B, achieving better performance than GPT-3.5 on all representative benchmarks for RTL code generation. Especially, it outperforms GPT-4 in VerilogEval Machine benchmark. This remarkable balance between accuracy and efficiency is made possible by leveraging our new RTL code dataset and a customized LLM algorithm, both of which have been made fully open-source.
format Preprint
id arxiv_https___arxiv_org_abs_2312_08617
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution
Liu, Shang
Fang, Wenji
Lu, Yao
Zhang, Qijun
Zhang, Hongce
Xie, Zhiyao
Programming Languages
Hardware Architecture
The automatic generation of RTL code (e.g., Verilog) using natural language instructions and large language models (LLMs) has attracted significant research interest recently. However, most existing approaches heavily rely on commercial LLMs such as ChatGPT, while open-source LLMs tailored for this specific design generation task exhibit notably inferior performance. The absence of high-quality open-source solutions restricts the flexibility and data privacy of this emerging technique. In this study, we present a new customized LLM solution with a modest parameter count of only 7B, achieving better performance than GPT-3.5 on all representative benchmarks for RTL code generation. Especially, it outperforms GPT-4 in VerilogEval Machine benchmark. This remarkable balance between accuracy and efficiency is made possible by leveraging our new RTL code dataset and a customized LLM algorithm, both of which have been made fully open-source.
title RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution
topic Programming Languages
Hardware Architecture
url https://arxiv.org/abs/2312.08617