HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Qirui, Shuai, Jingxian, Chen, Shuangwu, Ye, Shenghao, Wen, Zijian, Su, Xufei, Jin, Jie, Li, Jiangming, Chen, Jun, Tan, Xiaobin, Yang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SQLForge: Synthesizing Reliable and Diverse Data to Enhance Text-to-SQL Reasoning in LLMs
von: Guo, Yu, et al.
Veröffentlicht: (2025)
von: Guo, Yu, et al.
Veröffentlicht: (2025)
Rethinking Table Pruning in TableQA: From Sequential Revisions to Gold Trajectory-Supervised Parallel Search
von: Guo, Yu, et al.
Veröffentlicht: (2026)
von: Guo, Yu, et al.
Veröffentlicht: (2026)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
Rubric-Guided Process Reward for Stepwise Model Routing
von: Ye, Shenghao, et al.
Veröffentlicht: (2026)
von: Ye, Shenghao, et al.
Veröffentlicht: (2026)
Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspective
von: Ye, Shenghao, et al.
Veröffentlicht: (2026)
von: Ye, Shenghao, et al.
Veröffentlicht: (2026)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2024)
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2024)
SecCodeBench-V2 Technical Report
von: Chen, Longfei, et al.
Veröffentlicht: (2026)
von: Chen, Longfei, et al.
Veröffentlicht: (2026)
CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs
von: Li, Shiyang, et al.
Veröffentlicht: (2026)
von: Li, Shiyang, et al.
Veröffentlicht: (2026)
SecIC3: Customizing IC3 for Hardware Security Verification
von: Tan, Qinhan, et al.
Veröffentlicht: (2026)
von: Tan, Qinhan, et al.
Veröffentlicht: (2026)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
von: Wang, Yanlin, et al.
Veröffentlicht: (2026)
von: Wang, Yanlin, et al.
Veröffentlicht: (2026)
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
von: Jing, Pengfei, et al.
Veröffentlicht: (2024)
von: Jing, Pengfei, et al.
Veröffentlicht: (2024)
HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
von: Zhu, Qihui, et al.
Veröffentlicht: (2026)
von: Zhu, Qihui, et al.
Veröffentlicht: (2026)
When TableQA Meets Noise: A Dual Denoising Framework for Complex Questions and Large-scale Tables
von: Ye, Shenghao, et al.
Veröffentlicht: (2025)
von: Ye, Shenghao, et al.
Veröffentlicht: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
von: Zhang, Jiawen, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawen, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
SecReEvalBench: A Multi-turned Security Resilience Evaluation Benchmark for Large Language Models
von: Cui, Huining, et al.
Veröffentlicht: (2025)
von: Cui, Huining, et al.
Veröffentlicht: (2025)
SecCodePRM: A Process Reward Model for Code Security
von: Yu, Weichen, et al.
Veröffentlicht: (2026)
von: Yu, Weichen, et al.
Veröffentlicht: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
von: Yuan, He Yang, et al.
Veröffentlicht: (2026)
von: Yuan, He Yang, et al.
Veröffentlicht: (2026)
LiveSecBench: A Dynamic and Event-Driven Safety Benchmark for Chinese Language Model Applications
von: Li, Yudong, et al.
Veröffentlicht: (2025)
von: Li, Yudong, et al.
Veröffentlicht: (2025)
CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
von: Wahréus, Johan, et al.
Veröffentlicht: (2025)
von: Wahréus, Johan, et al.
Veröffentlicht: (2025)
SecFSM: Knowledge Graph-Guided Verilog Code Generation for Secure Finite State Machines in Systems-on-Chip
von: Hu, Ziteng, et al.
Veröffentlicht: (2025)
von: Hu, Ziteng, et al.
Veröffentlicht: (2025)
SecCoder: Towards Generalizable and Robust Secure Code Generation
von: Zhang, Boyu, et al.
Veröffentlicht: (2024)
von: Zhang, Boyu, et al.
Veröffentlicht: (2024)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
von: Zhao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Zhao, Yuxuan, et al.
Veröffentlicht: (2026)
HW-NAS-Bench:Hardware-Aware Neural Architecture Search Benchmark
von: Li, Chaojian, et al.
Veröffentlicht: (2021)
von: Li, Chaojian, et al.
Veröffentlicht: (2021)
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
von: Sukthanker, Rhea Sanjay, et al.
Veröffentlicht: (2024)
von: Sukthanker, Rhea Sanjay, et al.
Veröffentlicht: (2024)
CellularSpecSec-Bench: A Staged Benchmark for Evidence-Grounded Interpretation and Security Reasoning over 3GPP Specifications
von: Xie, Ke, et al.
Veröffentlicht: (2026)
von: Xie, Ke, et al.
Veröffentlicht: (2026)
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
von: Shi, Shaojie, et al.
Veröffentlicht: (2026)
von: Shi, Shaojie, et al.
Veröffentlicht: (2026)
CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
von: Zhou, Yu, et al.
Veröffentlicht: (2024)
von: Zhou, Yu, et al.
Veröffentlicht: (2024)
MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2026)
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2026)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SQLForge: Synthesizing Reliable and Diverse Data to Enhance Text-to-SQL Reasoning in LLMs
von: Guo, Yu, et al.
Veröffentlicht: (2025) -
Rethinking Table Pruning in TableQA: From Sequential Revisions to Gold Trajectory-Supervised Parallel Search
von: Guo, Yu, et al.
Veröffentlicht: (2026) -
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
von: Shen, Chihao, et al.
Veröffentlicht: (2025) -
Rubric-Guided Process Reward for Stepwise Model Routing
von: Ye, Shenghao, et al.
Veröffentlicht: (2026) -
Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspective
von: Ye, Shenghao, et al.
Veröffentlicht: (2026)