SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Qi, Tang, Yifeng, Wang, Qinghua, Zhao, Lanyang, Zhang, Pengji, Qing, Yuhao, Yao, Xin, Huang, Dong, Zhang, Lin, Ji, Zhuoran |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
Reflection-Driven Control for Trustworthy Code Agents
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
MCGMark: An Encodable and Robust Online Watermark for Tracing LLM-Generated Malicious Code
by: Ning, Kaiwen, et al.
Published: (2024)
by: Ning, Kaiwen, et al.
Published: (2024)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
by: Miculicich, Lesly, et al.
Published: (2025)
by: Miculicich, Lesly, et al.
Published: (2025)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
Exploiting LLM Agent Supply Chains via Payload-less Skills
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
Verbatim Data Transcription Failures in LLM Code Generation: A State-Tracking Stress Test
by: Haque, Mohd Ariful, et al.
Published: (2026)
by: Haque, Mohd Ariful, et al.
Published: (2026)
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026)
by: Xuan, Zhou, et al.
Published: (2026)
Jailbreak Distillation: Renewable Safety Benchmarking
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
LLM Security Guard for Code
by: Kavian, Arya, et al.
Published: (2024)
by: Kavian, Arya, et al.
Published: (2024)
ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents
by: Wu, Jiangrong, et al.
Published: (2026)
by: Wu, Jiangrong, et al.
Published: (2026)
MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
by: Tan, Zhuoran, et al.
Published: (2026)
by: Tan, Zhuoran, et al.
Published: (2026)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
ORCAS: Obfuscation-Resilient Binary Code Similarity Analysis using Dominance Enhanced Semantic Graph
by: Wang, Yufeng, et al.
Published: (2025)
by: Wang, Yufeng, et al.
Published: (2025)
How to Compare the Security of Code Written by Humans to LLM-generated Code
by: Balebako, Rebecca, et al.
Published: (2026)
by: Balebako, Rebecca, et al.
Published: (2026)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
KEENHash: Hashing Programs into Function-Aware Embeddings for Large-Scale Binary Code Similarity Analysis
by: Liu, Zhijie, et al.
Published: (2025)
by: Liu, Zhijie, et al.
Published: (2025)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026)
by: Ge, Yifei, et al.
Published: (2026)
NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist
by: Bertram, Johannes, et al.
Published: (2026)
by: Bertram, Johannes, et al.
Published: (2026)
CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation
by: Manuel, Dylan, et al.
Published: (2025)
by: Manuel, Dylan, et al.
Published: (2025)
HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Security Weaknesses of Copilot-Generated Code in GitHub Projects: An Empirical Study
by: Fu, Yujia, et al.
Published: (2023)
by: Fu, Yujia, et al.
Published: (2023)
Finding Privacy-relevant Source Code
by: Tang, Feiyang, et al.
Published: (2024)
by: Tang, Feiyang, et al.
Published: (2024)
Beyond Imprecise Distance Metrics: Trace-Guided Directed Greybox Fuzzing via LLM-Predicted Call Stacks
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
SAFuzz: Semantic-Guided Adaptive Fuzzing for LLM-Generated Code
by: Yang, Ziyi, et al.
Published: (2026)
by: Yang, Ziyi, et al.
Published: (2026)
An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code
by: Elsayed, Mohamed, et al.
Published: (2026)
by: Elsayed, Mohamed, et al.
Published: (2026)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks?
by: Zhang, Ying, et al.
Published: (2023)
by: Zhang, Ying, et al.
Published: (2023)
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
by: Xu, Hanxiang, et al.
Published: (2024)
by: Xu, Hanxiang, et al.
Published: (2024)
Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences
by: Hasan, Mohammad Saqib, et al.
Published: (2025)
by: Hasan, Mohammad Saqib, et al.
Published: (2025)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
CHASE: LLM Agents for Dissecting Malicious PyPI Packages
by: Toda, Takaaki, et al.
Published: (2026)
by: Toda, Takaaki, et al.
Published: (2026)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
by: Liu, Bin, et al.
Published: (2025)
by: Liu, Bin, et al.
Published: (2025)
SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
Similar Items
-
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026) -
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026) -
Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements
by: Li, Yue, et al.
Published: (2026) -
Reflection-Driven Control for Trustworthy Code Agents
by: Wang, Bin, et al.
Published: (2025) -
MCGMark: An Encodable and Robust Online Watermark for Tracing LLM-Generated Malicious Code
by: Ning, Kaiwen, et al.
Published: (2024)