AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Ruidi, Zhang, Grace Li, Drechsler, Rolf, Schlichtmann, Ulf, Li, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CorrectBench: Automatic Testbench Generation with Functional Self-Correction using LLMs for HDL Design
by: Qiu, Ruidi, et al.
Published: (2024)
by: Qiu, Ruidi, et al.
Published: (2024)
CorrectHDL: Agentic HDL Design with LLMs Leveraging High-Level Synthesis as Reference
by: Xu, Kangwei, et al.
Published: (2025)
by: Xu, Kangwei, et al.
Published: (2025)
VFocus: Better Verilog Generation from Large Language Model via Focused Reasoning
by: Zhao, Zhuorui, et al.
Published: (2025)
by: Zhao, Zhuorui, et al.
Published: (2025)
Paradigm-Based Automatic HDL Code Generation Using LLMs
by: Sun, Wenhao, et al.
Published: (2025)
by: Sun, Wenhao, et al.
Published: (2025)
HLSTester: Efficient Testing of Behavioral Discrepancies with LLMs for High-Level Synthesis
by: Xu, Kangwei, et al.
Published: (2025)
by: Xu, Kangwei, et al.
Published: (2025)
Accurate and Extensible Symbolic Execution of Binary Code based on Formal ISA Semantics
by: Tempel, Sören, et al.
Published: (2024)
by: Tempel, Sören, et al.
Published: (2024)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
AutoMCQ -- Automatically Generate Code Comprehension Questions using GenAI
by: Goodfellow, Martin, et al.
Published: (2025)
by: Goodfellow, Martin, et al.
Published: (2025)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
by: Chou, Jason, et al.
Published: (2025)
by: Chou, Jason, et al.
Published: (2025)
VRank: Enhancing Verilog Code Generation from Large Language Models via Self-Consistency
by: Zhao, Zhuorui, et al.
Published: (2025)
by: Zhao, Zhuorui, et al.
Published: (2025)
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
iScript: A Domain-Adapted Large Language Model and Benchmark for Physical Design Tcl Script Generation
by: Xu, Ning, et al.
Published: (2026)
by: Xu, Ning, et al.
Published: (2026)
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
by: Cai, Jianfeng, et al.
Published: (2026)
by: Cai, Jianfeng, et al.
Published: (2026)
The CodeInverter Suite: Control-Flow and Data-Mapping Augmented Binary Decompilation with LLMs
by: Liu, Peipei, et al.
Published: (2025)
by: Liu, Peipei, et al.
Published: (2025)
Bench4HLS: End-to-End Evaluation of LLMs in High-Level Synthesis Code Generation
by: Khan, M Zafir Sadik, et al.
Published: (2026)
by: Khan, M Zafir Sadik, et al.
Published: (2026)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
by: Yang, Shu, et al.
Published: (2026)
by: Yang, Shu, et al.
Published: (2026)
Understanding and Detecting Platform-Specific Violations in Android Auto Apps
by: Fakorede, Moshood, et al.
Published: (2025)
by: Fakorede, Moshood, et al.
Published: (2025)
Phaedrus: Predicting Dynamic Application Behavior with Lightweight Generative Models and LLMs
by: Chatterjee, Bodhisatwa, et al.
Published: (2024)
by: Chatterjee, Bodhisatwa, et al.
Published: (2024)
ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch
by: Si, Yuan, et al.
Published: (2025)
by: Si, Yuan, et al.
Published: (2025)
Automated Discovery of Test Oracles for Database Management Systems Using LLMs
by: Mang, Qiuyang, et al.
Published: (2025)
by: Mang, Qiuyang, et al.
Published: (2025)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
by: Dong, Honghua, et al.
Published: (2025)
by: Dong, Honghua, et al.
Published: (2025)
Strengthening Programming Comprehension in Large Language Models through Code Generation
by: Ren, Xiaoning, et al.
Published: (2025)
by: Ren, Xiaoning, et al.
Published: (2025)
MLIR-Smith: A Novel Random Program Generator for Evaluating Compiler Pipelines
by: Ates, Berke, et al.
Published: (2026)
by: Ates, Berke, et al.
Published: (2026)
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
by: Wang, Ming, et al.
Published: (2024)
by: Wang, Ming, et al.
Published: (2024)
Fully Automated Generation of Combinatorial Optimisation Systems Using Large Language Models
by: Karapetyan, Daniel
Published: (2025)
by: Karapetyan, Daniel
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Enabling Memory Safety of C Programs using LLMs
by: Mohammed, Nausheen, et al.
Published: (2024)
by: Mohammed, Nausheen, et al.
Published: (2024)
Python Symbolic Execution with LLM-powered Code Generation
by: Wang, Wenhan, et al.
Published: (2024)
by: Wang, Wenhan, et al.
Published: (2024)
DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Generating Equivalent Representations of Code By A Self-Reflection Approach
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
by: Peng, Qiwei, et al.
Published: (2024)
by: Peng, Qiwei, et al.
Published: (2024)
Specification-Guided Repair of Arithmetic Errors in Dafny Programs using LLMs
by: Wu, Valentina, et al.
Published: (2025)
by: Wu, Valentina, et al.
Published: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025)
by: Le-Cong, Thanh, et al.
Published: (2025)
SAT-DIFF: A Tree Diffing Framework Using SAT Solving
by: Geng, Chuqin, et al.
Published: (2024)
by: Geng, Chuqin, et al.
Published: (2024)
JustinANN: Realistic Test Generation for Java Programs Driven by Annotations
by: Cui, Baoquan, et al.
Published: (2025)
by: Cui, Baoquan, et al.
Published: (2025)
Reverse Chain: A Generic-Rule for LLMs to Master Multi-API Planning
by: Zhang, Yinger, et al.
Published: (2023)
by: Zhang, Yinger, et al.
Published: (2023)
LLMs Lean on Priors, Not Programming Language Semantics
by: Thimmaiah, Aditya, et al.
Published: (2025)
by: Thimmaiah, Aditya, et al.
Published: (2025)
Similar Items
-
CorrectBench: Automatic Testbench Generation with Functional Self-Correction using LLMs for HDL Design
by: Qiu, Ruidi, et al.
Published: (2024) -
CorrectHDL: Agentic HDL Design with LLMs Leveraging High-Level Synthesis as Reference
by: Xu, Kangwei, et al.
Published: (2025) -
VFocus: Better Verilog Generation from Large Language Model via Focused Reasoning
by: Zhao, Zhuorui, et al.
Published: (2025) -
Paradigm-Based Automatic HDL Code Generation Using LLMs
by: Sun, Wenhao, et al.
Published: (2025) -
HLSTester: Efficient Testing of Behavioral Discrepancies with LLMs for High-Level Synthesis
by: Xu, Kangwei, et al.
Published: (2025)