CorrectBench: Automatic Testbench Generation with Functional Self-Correction using LLMs for HDL Design
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Ruidi, Zhang, Grace Li, Drechsler, Rolf, Schlichtmann, Ulf, Li, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
by: Qiu, Ruidi, et al.
Published: (2024)
by: Qiu, Ruidi, et al.
Published: (2024)
CorrectHDL: Agentic HDL Design with LLMs Leveraging High-Level Synthesis as Reference
by: Xu, Kangwei, et al.
Published: (2025)
by: Xu, Kangwei, et al.
Published: (2025)
HLSTester: Efficient Testing of Behavioral Discrepancies with LLMs for High-Level Synthesis
by: Xu, Kangwei, et al.
Published: (2025)
by: Xu, Kangwei, et al.
Published: (2025)
VFocus: Better Verilog Generation from Large Language Model via Focused Reasoning
by: Zhao, Zhuorui, et al.
Published: (2025)
by: Zhao, Zhuorui, et al.
Published: (2025)
Log Parsing using LLMs with Self-Generated In-Context Learning and Self-Correction
by: Wu, Yifan, et al.
Published: (2024)
by: Wu, Yifan, et al.
Published: (2024)
LLM-Assisted Tool for Joint Generation of Formulas and Functions in Rule-Based Verification of Map Transformations
by: He, Ruidi, et al.
Published: (2025)
by: He, Ruidi, et al.
Published: (2025)
Classification-Based Automatic HDL Code Generation Using LLMs
by: Sun, Wenhao, et al.
Published: (2024)
by: Sun, Wenhao, et al.
Published: (2024)
Paradigm-Based Automatic HDL Code Generation Using LLMs
by: Sun, Wenhao, et al.
Published: (2025)
by: Sun, Wenhao, et al.
Published: (2025)
DesignCoder: Hierarchy-Aware and Self-Correcting UI Code Generation with Large Language Models
by: Chen, Yunnong, et al.
Published: (2025)
by: Chen, Yunnong, et al.
Published: (2025)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
by: Zhu, Hongda, et al.
Published: (2025)
by: Zhu, Hongda, et al.
Published: (2025)
Beyond Functional Correctness: Design Issues in AI IDE-Generated Large-Scale Projects
by: Kashif, Syed Mohammad, et al.
Published: (2026)
by: Kashif, Syed Mohammad, et al.
Published: (2026)
TOGLL: Correct and Strong Test Oracle Generation with LLMs
by: Hossain, Soneya Binta, et al.
Published: (2024)
by: Hossain, Soneya Binta, et al.
Published: (2024)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
Statistical Confidence in Functional Correctness: An Approach for AI Product Functional Correctness Evaluation
by: Albertini, Wallace, et al.
Published: (2026)
by: Albertini, Wallace, et al.
Published: (2026)
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
by: Peng, Yibo, et al.
Published: (2025)
by: Peng, Yibo, et al.
Published: (2025)
PatchZero: Zero-Shot Automatic Patch Correctness Assessment
by: Zhou, Xin, et al.
Published: (2023)
by: Zhou, Xin, et al.
Published: (2023)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
Lightweight Model Editing for LLMs to Correct Deprecated API Recommendations
by: Lin, Guancheng, et al.
Published: (2025)
by: Lin, Guancheng, et al.
Published: (2025)
AL-Bench: A Benchmark for Automatic Logging
by: Tan, Boyin, et al.
Published: (2025)
by: Tan, Boyin, et al.
Published: (2025)
FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
by: Zhou, Zhiping, et al.
Published: (2025)
by: Zhou, Zhiping, et al.
Published: (2025)
AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree Search
by: Li, Qingyao, et al.
Published: (2026)
by: Li, Qingyao, et al.
Published: (2026)
Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study
by: Santana Jr, E. G., et al.
Published: (2025)
by: Santana Jr, E. G., et al.
Published: (2025)
Are LLMs Correctly Integrated into Software Systems?
by: Shao, Yuchen, et al.
Published: (2024)
by: Shao, Yuchen, et al.
Published: (2024)
LLM-based Behaviour Driven Development for Hardware Design
by: Drechsler, Rolf, et al.
Published: (2025)
by: Drechsler, Rolf, et al.
Published: (2025)
Correctness Witnesses with Function Contracts
by: Heizmann, Matthias, et al.
Published: (2025)
by: Heizmann, Matthias, et al.
Published: (2025)
Ensuring Functional Correctness of Large Code Models with Selective Generation
by: Jeong, Jaewoo, et al.
Published: (2025)
by: Jeong, Jaewoo, et al.
Published: (2025)
SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated Code
by: Wang, Qinglin, et al.
Published: (2025)
by: Wang, Qinglin, et al.
Published: (2025)
CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
A Novel HDL Code Generator for Effectively Testing FPGA Logic Synthesis Compilers
by: Xu, Zhihao, et al.
Published: (2024)
by: Xu, Zhihao, et al.
Published: (2024)
On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
by: Naik, Atharva
Published: (2024)
by: Naik, Atharva
Published: (2024)
In-Context Learning as an Effective Estimator of Functional Correctness of LLM-Generated Code
by: Das, Susmita, et al.
Published: (2025)
by: Das, Susmita, et al.
Published: (2025)
Ensemble-Based Uncertainty Estimation for Code Correctness Estimation
by: Wei, Yunxiang, et al.
Published: (2026)
by: Wei, Yunxiang, et al.
Published: (2026)
CoCoST: Automatic Complex Code Generation with Online Searching and Correctness Testing
by: He, Xinyi, et al.
Published: (2024)
by: He, Xinyi, et al.
Published: (2024)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
by: Chou, Jason, et al.
Published: (2025)
by: Chou, Jason, et al.
Published: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
by: Zhang, Ruiyi, et al.
Published: (2026)
by: Zhang, Ruiyi, et al.
Published: (2026)
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
by: Allamanis, Miltiadis, et al.
Published: (2024)
by: Allamanis, Miltiadis, et al.
Published: (2024)
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
by: Li, Haiyang
Published: (2025)
by: Li, Haiyang
Published: (2025)
JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
by: Wang, Yiran, et al.
Published: (2025)
by: Wang, Yiran, et al.
Published: (2025)
Automatic Generation of Formal Specification and Verification Annotations Using LLMs and Test Oracles
by: Faria, João Pascoal, et al.
Published: (2026)
by: Faria, João Pascoal, et al.
Published: (2026)
Similar Items
-
AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
by: Qiu, Ruidi, et al.
Published: (2024) -
CorrectHDL: Agentic HDL Design with LLMs Leveraging High-Level Synthesis as Reference
by: Xu, Kangwei, et al.
Published: (2025) -
HLSTester: Efficient Testing of Behavioral Discrepancies with LLMs for High-Level Synthesis
by: Xu, Kangwei, et al.
Published: (2025) -
VFocus: Better Verilog Generation from Large Language Model via Focused Reasoning
by: Zhao, Zhuorui, et al.
Published: (2025) -
Log Parsing using LLMs with Self-Generated In-Context Learning and Self-Correction
by: Wu, Yifan, et al.
Published: (2024)