Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Junjie, Shen, Xinyue, Leng, Ye, Backes, Michael, Shen, Yun, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
by: Zhang, Yage, et al.
Published: (2026)
by: Zhang, Yage, et al.
Published: (2026)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026)
by: Qin, Kaihua, et al.
Published: (2026)
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models
by: Bruni, Marc, et al.
Published: (2025)
by: Bruni, Marc, et al.
Published: (2025)
DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation
by: Pathak, Abhijeet, et al.
Published: (2025)
by: Pathak, Abhijeet, et al.
Published: (2025)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
by: Dubniczky, Richard A., et al.
Published: (2025)
by: Dubniczky, Richard A., et al.
Published: (2025)
Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users
by: Bhardwaj, Arth, et al.
Published: (2026)
by: Bhardwaj, Arth, et al.
Published: (2026)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
by: Tang, Yuheng, et al.
Published: (2026)
by: Tang, Yuheng, et al.
Published: (2026)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study
by: Tamberg, Karl, et al.
Published: (2024)
by: Tamberg, Karl, et al.
Published: (2024)
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Repository-Level Graph Representation Learning for Enhanced Security Patch Detection
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
by: Shen, Qingchao, et al.
Published: (2026)
by: Shen, Qingchao, et al.
Published: (2026)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
by: Chu, Zhaoyang, et al.
Published: (2025)
by: Chu, Zhaoyang, et al.
Published: (2025)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
by: Hu, Junze, et al.
Published: (2025)
by: Hu, Junze, et al.
Published: (2025)
Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
by: Patir, Rupam, et al.
Published: (2025)
by: Patir, Rupam, et al.
Published: (2025)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
by: Mitropoulos, Dimitris, et al.
Published: (2026)
by: Mitropoulos, Dimitris, et al.
Published: (2026)
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
by: Xie, Xuan, et al.
Published: (2024)
by: Xie, Xuan, et al.
Published: (2024)
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
by: Jiang, Weipeng, et al.
Published: (2026)
by: Jiang, Weipeng, et al.
Published: (2026)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist
by: Bertram, Johannes, et al.
Published: (2026)
by: Bertram, Johannes, et al.
Published: (2026)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
LLM-enabled Applications Require System-Level Threat Monitoring
by: Zhang, Yedi, et al.
Published: (2026)
by: Zhang, Yedi, et al.
Published: (2026)
SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability Detection
by: Yu, Lei, et al.
Published: (2025)
by: Yu, Lei, et al.
Published: (2025)
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
by: Lu, Xiaoyue, et al.
Published: (2026)
by: Lu, Xiaoyue, et al.
Published: (2026)
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
by: Bommarito II, Michael J.
Published: (2026)
by: Bommarito II, Michael J.
Published: (2026)
Jailbreak Distillation: Renewable Safety Benchmarking
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
Smart-LLaMA: Two-Stage Post-Training of Large Language Models for Smart Contract Vulnerability Detection and Explanation
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
Towards Privacy-Preserving Code Generation: Differentially Private Code Language Models
by: Catal, Melih, et al.
Published: (2025)
by: Catal, Melih, et al.
Published: (2025)
SecCodeBench-V2 Technical Report
by: Chen, Longfei, et al.
Published: (2026)
by: Chen, Longfei, et al.
Published: (2026)
Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability Detection
by: Yu, Lei, et al.
Published: (2025)
by: Yu, Lei, et al.
Published: (2025)
Similar Items
-
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
by: Zhang, Yage, et al.
Published: (2026) -
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025) -
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026) -
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026) -
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models
by: Bruni, Marc, et al.
Published: (2025)