MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Steinberg, Jonathan, Gal, Oren |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024)
by: Mahyari, Andrew A
Published: (2024)
SecCodeBench-V2 Technical Report
by: Chen, Longfei, et al.
Published: (2026)
by: Chen, Longfei, et al.
Published: (2026)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
by: Hu, Junze, et al.
Published: (2025)
by: Hu, Junze, et al.
Published: (2025)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
by: Wang, Su, et al.
Published: (2026)
by: Wang, Su, et al.
Published: (2026)
MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
by: Storhaug, André, et al.
Published: (2026)
by: Storhaug, André, et al.
Published: (2026)
From Detection to Prevention: Explaining Security-Critical Code to Avoid Vulnerabilities
by: Krishnamurthy, Ranjith, et al.
Published: (2026)
by: Krishnamurthy, Ranjith, et al.
Published: (2026)
Code Security Vulnerability Repair Using Reinforcement Learning with Large Language Models
by: Islam, Nafis Tanveer, et al.
Published: (2024)
by: Islam, Nafis Tanveer, et al.
Published: (2024)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code
by: Blain, Dominik, et al.
Published: (2026)
by: Blain, Dominik, et al.
Published: (2026)
CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming Solutions
by: Shi, Jingwei, et al.
Published: (2026)
by: Shi, Jingwei, et al.
Published: (2026)
SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents
by: Begimher, Daniel, et al.
Published: (2026)
by: Begimher, Daniel, et al.
Published: (2026)
MAVUL: Multi-Agent Vulnerability Detection via Contextual Reasoning and Interactive Refinement
by: Li, Youpeng, et al.
Published: (2025)
by: Li, Youpeng, et al.
Published: (2025)
Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning
by: Yan, Shenao, et al.
Published: (2026)
by: Yan, Shenao, et al.
Published: (2026)
Reflection-Driven Control for Trustworthy Code Agents
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning
by: Ni, Ronghao, et al.
Published: (2026)
by: Ni, Ronghao, et al.
Published: (2026)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
by: Mitropoulos, Dimitris, et al.
Published: (2026)
by: Mitropoulos, Dimitris, et al.
Published: (2026)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
An Empirical Study of Vulnerabilities in Python Packages and Their Detection
by: Quan, Haowei, et al.
Published: (2025)
by: Quan, Haowei, et al.
Published: (2025)
Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
LLbezpeky: Leveraging Large Language Models for Vulnerability Detection
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
An Empirical Study of Vulnerability Detection using Federated Learning
by: Zhou, Peiheng, et al.
Published: (2024)
by: Zhou, Peiheng, et al.
Published: (2024)
How Effective Are Neural Networks for Fixing Security Vulnerabilities
by: Wu, Yi, et al.
Published: (2023)
by: Wu, Yi, et al.
Published: (2023)
Detecting Vulnerabilities from Issue Reports for Internet-of-Things
by: Masoumzadeh, Sogol
Published: (2025)
by: Masoumzadeh, Sogol
Published: (2025)
Semantic Denial of Service in LLM-controlled robots
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025)
by: Garg, Aayush, et al.
Published: (2025)
Graph Neural Networks for Vulnerability Detection: A Counterfactual Explanation
by: Chu, Zhaoyang, et al.
Published: (2024)
by: Chu, Zhaoyang, et al.
Published: (2024)
Needles at Scale: LLM-Assisted Target Selection for Windows Vulnerability Research
by: Bommarito II, Michael J.
Published: (2026)
by: Bommarito II, Michael J.
Published: (2026)
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization
by: Kong, Ziqiao, et al.
Published: (2026)
by: Kong, Ziqiao, et al.
Published: (2026)
VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization
by: Li, Youpeng, et al.
Published: (2025)
by: Li, Youpeng, et al.
Published: (2025)
Focus on What Matters: Fisher-Guided Adaptive Multimodal Fusion for Vulnerability Detection
by: Bian, Yun, et al.
Published: (2026)
by: Bian, Yun, et al.
Published: (2026)
A Comprehensive Study of Exploitable Patterns in Smart Contracts: From Vulnerability to Defense
by: Ding, Yuchen, et al.
Published: (2025)
by: Ding, Yuchen, et al.
Published: (2025)
Bridging Semantics & Structure for Software Vulnerability Detection using Hybrid Network Models
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Data and Context Matter: Towards Generalizing AI-based Software Vulnerability Detection
by: Safdar, Rijha, et al.
Published: (2025)
by: Safdar, Rijha, et al.
Published: (2025)
GPTScan: Detecting Logic Vulnerabilities in Smart Contracts by Combining GPT with Program Analysis
by: Sun, Yuqiang, et al.
Published: (2023)
by: Sun, Yuqiang, et al.
Published: (2023)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2026)
by: Li, Youpeng, et al.
Published: (2026)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
Similar Items
-
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025) -
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
by: Liu, Yu, et al.
Published: (2024) -
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024) -
SecCodeBench-V2 Technical Report
by: Chen, Longfei, et al.
Published: (2026) -
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
by: Hu, Junze, et al.
Published: (2025)