Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security
Fuente:
arXiv
Salvato in:
| Autore principale: | Chua, Gabriel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
di: Li, Xinghang, et al.
Pubblicazione: (2025)
di: Li, Xinghang, et al.
Pubblicazione: (2025)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
di: Shen, Chihao, et al.
Pubblicazione: (2025)
di: Shen, Chihao, et al.
Pubblicazione: (2025)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
di: Chen, Qirui, et al.
Pubblicazione: (2026)
di: Chen, Qirui, et al.
Pubblicazione: (2026)
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
di: Happe, Andreas, et al.
Pubblicazione: (2025)
di: Happe, Andreas, et al.
Pubblicazione: (2025)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
di: Fu, Yuchuan, et al.
Pubblicazione: (2025)
di: Fu, Yuchuan, et al.
Pubblicazione: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
di: Zhang, Hanrong, et al.
Pubblicazione: (2024)
di: Zhang, Hanrong, et al.
Pubblicazione: (2024)
DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation
di: Pathak, Abhijeet, et al.
Pubblicazione: (2025)
di: Pathak, Abhijeet, et al.
Pubblicazione: (2025)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
di: Shao, Minghao, et al.
Pubblicazione: (2025)
di: Shao, Minghao, et al.
Pubblicazione: (2025)
Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
LLM-CSEC: Empirical Evaluation of Security in C/C++ Code Generated by Large Language Models
di: Shahid, Muhammad Usman, et al.
Pubblicazione: (2025)
di: Shahid, Muhammad Usman, et al.
Pubblicazione: (2025)
When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents
di: Kozak, Matous, et al.
Pubblicazione: (2025)
di: Kozak, Matous, et al.
Pubblicazione: (2025)
Security of LLM-generated Code: A Comparative Analysis
di: Morkonda, Srivathsan G, et al.
Pubblicazione: (2026)
di: Morkonda, Srivathsan G, et al.
Pubblicazione: (2026)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
di: Zhang, Dongsen, et al.
Pubblicazione: (2025)
di: Zhang, Dongsen, et al.
Pubblicazione: (2025)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
di: Yuan, He Yang, et al.
Pubblicazione: (2026)
di: Yuan, He Yang, et al.
Pubblicazione: (2026)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
A Mixture of Linear Corrections Generates Secure Code
di: Yu, Weichen, et al.
Pubblicazione: (2025)
di: Yu, Weichen, et al.
Pubblicazione: (2025)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
di: Nie, Yuzhou, et al.
Pubblicazione: (2024)
di: Nie, Yuzhou, et al.
Pubblicazione: (2024)
On the (In)Security of LLM App Stores
di: Hou, Xinyi, et al.
Pubblicazione: (2024)
di: Hou, Xinyi, et al.
Pubblicazione: (2024)
Benchmarking LLM-Based Static Analysis for Secure Smart Contract Development: Reliability, Limitations, and Potential Hybrid Solutions
di: Susan, Stefan-Claudiu, et al.
Pubblicazione: (2026)
di: Susan, Stefan-Claudiu, et al.
Pubblicazione: (2026)
A Framework for Formalizing LLM Agent Security
di: Siu, Vincent, et al.
Pubblicazione: (2026)
di: Siu, Vincent, et al.
Pubblicazione: (2026)
Enhancing Reliability in LLM-Based Secure Code Generation
di: Kharma, Mohammed F., et al.
Pubblicazione: (2026)
di: Kharma, Mohammed F., et al.
Pubblicazione: (2026)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
di: Chu, Junjie, et al.
Pubblicazione: (2026)
di: Chu, Junjie, et al.
Pubblicazione: (2026)
Ocassionally Secure: A Comparative Analysis of Code Generation Assistants
di: Elgedawy, Ran, et al.
Pubblicazione: (2024)
di: Elgedawy, Ran, et al.
Pubblicazione: (2024)
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
Information Security Based on LLM Approaches: A Review
di: Gong, Chang, et al.
Pubblicazione: (2025)
di: Gong, Chang, et al.
Pubblicazione: (2025)
CodeBC: A More Secure Large Language Model for Smart Contract Code Generation in Blockchain
di: Wang, Lingxiang, et al.
Pubblicazione: (2025)
di: Wang, Lingxiang, et al.
Pubblicazione: (2025)
Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
di: Patir, Rupam, et al.
Pubblicazione: (2025)
di: Patir, Rupam, et al.
Pubblicazione: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
VibeGuard: A Security Gate Framework for AI-Generated Code
di: Xie, Ying
Pubblicazione: (2026)
di: Xie, Ying
Pubblicazione: (2026)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
di: Kim, Jungin, et al.
Pubblicazione: (2025)
di: Kim, Jungin, et al.
Pubblicazione: (2025)
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models
di: Bruni, Marc, et al.
Pubblicazione: (2025)
di: Bruni, Marc, et al.
Pubblicazione: (2025)
Artificial-Intelligence Generated Code Considered Harmful: A Road Map for Secure and High-Quality Code Generation
di: Chong, Chun Jie, et al.
Pubblicazione: (2024)
di: Chong, Chun Jie, et al.
Pubblicazione: (2024)
SkillTester: Benchmarking Utility and Security of Agent Skills
di: Wang, Leye, et al.
Pubblicazione: (2026)
di: Wang, Leye, et al.
Pubblicazione: (2026)
Security in LLM-as-a-Judge: A Comprehensive SoK
di: Masoud, Aiman Al, et al.
Pubblicazione: (2026)
di: Masoud, Aiman Al, et al.
Pubblicazione: (2026)
LLM Agents Should Employ Security Principles
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
aiXamine: Simplified LLM Safety and Security
di: Deniz, Fatih, et al.
Pubblicazione: (2025)
di: Deniz, Fatih, et al.
Pubblicazione: (2025)
ESAA-Security: An Event-Sourced, Verifiable Architecture for Agent-Assisted Security Audits of AI-Generated Code
di: Filho, Elzo Brito dos Santos
Pubblicazione: (2026)
di: Filho, Elzo Brito dos Santos
Pubblicazione: (2026)
Demo: SGCode: A Flexible Prompt-Optimizing System for Secure Generation of Code
di: Ton, Khiem, et al.
Pubblicazione: (2024)
di: Ton, Khiem, et al.
Pubblicazione: (2024)
Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
di: Mitropoulos, Dimitris, et al.
Pubblicazione: (2026)
di: Mitropoulos, Dimitris, et al.
Pubblicazione: (2026)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
di: Kharma, Mohammed, et al.
Pubblicazione: (2026)
di: Kharma, Mohammed, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
di: Li, Xinghang, et al.
Pubblicazione: (2025) -
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
di: Shen, Chihao, et al.
Pubblicazione: (2025) -
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
di: Chen, Qirui, et al.
Pubblicazione: (2026) -
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
di: Happe, Andreas, et al.
Pubblicazione: (2025) -
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
di: Fu, Yuchuan, et al.
Pubblicazione: (2025)