AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Yuchao, Fang, Wenji, Wang, Jing, Li, Wenkai, Guo, Ziyan, Xie, Zhiyao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911722671964160
author Wu, Yuchao
Fang, Wenji
Wang, Jing
Li, Wenkai
Guo, Ziyan
Xie, Zhiyao
author_facet Wu, Yuchao
Fang, Wenji
Wang, Jing
Li, Wenkai
Guo, Ziyan
Xie, Zhiyao
contents Assertion-based verification (ABV) is a cornerstone of modern hardware design, yet manually translating design intent into formal SystemVerilog Assertions (SVAs) remains labor-intensive and error-prone. While Large Language Models (LLMs) show promise for automating this process, existing benchmarks remain limited by unrealistic task formulations, weak specification inputs, and oversimplified evaluation. To address these limitations, we introduce AssertLLM2, an open-source benchmark for realistic assertion generation in hardware verification. AssertLLM2 contains 83 real-world designs across 13 functional categories. For each design, the benchmark provides a structured design specification, a verified dependency-complete golden RTL, and systematically mutated buggy RTL variants. These support two practical settings: bug-prevention, where assertions are generated from specifications to guard against design errors, and bug-hunting, where assertions are generated to expose discrepancies between intended behavior and faulty implementations. To the best of our knowledge, AssertLLM2 is the first benchmark to explicitly use buggy RTL as input to evaluate bug-detection capability. AssertLLM2 further adopts a more rigorous evaluation framework spanning syntactic validity, formal provability, coverage, and mutation-based bug detection. Our benchmark enables a more realistic and extensive assessment of assertion generation and establishes rigorous baselines for state-of-the-art LLMs in practical hardware verification.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27472
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications
Wu, Yuchao
Fang, Wenji
Wang, Jing
Li, Wenkai
Guo, Ziyan
Xie, Zhiyao
Hardware Architecture
Artificial Intelligence
Assertion-based verification (ABV) is a cornerstone of modern hardware design, yet manually translating design intent into formal SystemVerilog Assertions (SVAs) remains labor-intensive and error-prone. While Large Language Models (LLMs) show promise for automating this process, existing benchmarks remain limited by unrealistic task formulations, weak specification inputs, and oversimplified evaluation. To address these limitations, we introduce AssertLLM2, an open-source benchmark for realistic assertion generation in hardware verification. AssertLLM2 contains 83 real-world designs across 13 functional categories. For each design, the benchmark provides a structured design specification, a verified dependency-complete golden RTL, and systematically mutated buggy RTL variants. These support two practical settings: bug-prevention, where assertions are generated from specifications to guard against design errors, and bug-hunting, where assertions are generated to expose discrepancies between intended behavior and faulty implementations. To the best of our knowledge, AssertLLM2 is the first benchmark to explicitly use buggy RTL as input to evaluate bug-detection capability. AssertLLM2 further adopts a more rigorous evaluation framework spanning syntactic validity, formal provability, coverage, and mutation-based bug detection. Our benchmark enables a more realistic and extensive assessment of assertion generation and establishes rigorous baselines for state-of-the-art LLMs in practical hardware verification.
title AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications
topic Hardware Architecture
Artificial Intelligence
url https://arxiv.org/abs/2605.27472