Reflection-Driven Control for Trustworthy Code Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Bin, Quan, Jiazheng, Yu, Xingrui, Hu, Hansen, Yuhao, Tsang, Ivor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Generate Secure Code via Token-Level Rewards
by: Quan, Jiazheng, et al.
Published: (2026)
by: Quan, Jiazheng, et al.
Published: (2026)
Python Fuzzing for Trustworthy Machine Learning Frameworks
by: Yegorov, Ilya, et al.
Published: (2024)
by: Yegorov, Ilya, et al.
Published: (2024)
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
by: Sheh, Raymond K., et al.
Published: (2025)
by: Sheh, Raymond K., et al.
Published: (2025)
Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
by: Patir, Rupam, et al.
Published: (2025)
by: Patir, Rupam, et al.
Published: (2025)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
by: Hu, Junze, et al.
Published: (2025)
by: Hu, Junze, et al.
Published: (2025)
Options, Not Clicks: Lattice Refinement for Consent-Driven MCP Authorization
by: Li, Ying, et al.
Published: (2026)
by: Li, Ying, et al.
Published: (2026)
DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation
by: Pathak, Abhijeet, et al.
Published: (2025)
by: Pathak, Abhijeet, et al.
Published: (2025)
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
by: Tracy, Tyler, et al.
Published: (2026)
by: Tracy, Tyler, et al.
Published: (2026)
On the Security Risks of ML-based Malware Detection Systems: A Survey
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
by: Wang, Su, et al.
Published: (2026)
by: Wang, Su, et al.
Published: (2026)
Towards Privacy-Preserving Code Generation: Differentially Private Code Language Models
by: Catal, Melih, et al.
Published: (2025)
by: Catal, Melih, et al.
Published: (2025)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
by: Gu, Shanzhi, et al.
Published: (2025)
by: Gu, Shanzhi, et al.
Published: (2025)
MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
MALSIGHT: Exploring Malicious Source Code and Benign Pseudocode for Iterative Binary Malware Summarization
by: Lu, Haolang, et al.
Published: (2024)
by: Lu, Haolang, et al.
Published: (2024)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
PatUntrack: Automated Generating Patch Examples for Issue Reports without Tracked Insecure Code
by: Jiang, Ziyou, et al.
Published: (2024)
by: Jiang, Ziyou, et al.
Published: (2024)
SecCodeBench-V2 Technical Report
by: Chen, Longfei, et al.
Published: (2026)
by: Chen, Longfei, et al.
Published: (2026)
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Focus on What Matters: Fisher-Guided Adaptive Multimodal Fusion for Vulnerability Detection
by: Bian, Yun, et al.
Published: (2026)
by: Bian, Yun, et al.
Published: (2026)
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023)
by: Al-Kaswan, Ali, et al.
Published: (2023)
AI Code Generators for Security: Friend or Foe?
by: Natella, Roberto, et al.
Published: (2024)
by: Natella, Roberto, et al.
Published: (2024)
OpenSage: Self-programming Agent Generation Engine
by: Li, Hongwei, et al.
Published: (2026)
by: Li, Hongwei, et al.
Published: (2026)
Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning
by: Yan, Shenao, et al.
Published: (2026)
by: Yan, Shenao, et al.
Published: (2026)
An Empirical Study of Vulnerability Detection using Federated Learning
by: Zhou, Peiheng, et al.
Published: (2024)
by: Zhou, Peiheng, et al.
Published: (2024)
An Empirical Study of Vulnerabilities in Python Packages and Their Detection
by: Quan, Haowei, et al.
Published: (2025)
by: Quan, Haowei, et al.
Published: (2025)
Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
by: Chu, Zhaoyang, et al.
Published: (2025)
by: Chu, Zhaoyang, et al.
Published: (2025)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024)
by: Mahyari, Andrew A
Published: (2024)
MAVUL: Multi-Agent Vulnerability Detection via Contextual Reasoning and Interactive Refinement
by: Li, Youpeng, et al.
Published: (2025)
by: Li, Youpeng, et al.
Published: (2025)
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
by: Gagnon, Charles E., et al.
Published: (2025)
by: Gagnon, Charles E., et al.
Published: (2025)
An Empirical Study of Code Obfuscation Practices in the Google Play Store
by: Niroshan, Akila, et al.
Published: (2025)
by: Niroshan, Akila, et al.
Published: (2025)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Similar Items
-
Learning to Generate Secure Code via Token-Level Rewards
by: Quan, Jiazheng, et al.
Published: (2026) -
Python Fuzzing for Trustworthy Machine Learning Frameworks
by: Yegorov, Ilya, et al.
Published: (2024) -
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
by: Steinberg, Jonathan, et al.
Published: (2026) -
Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
by: Sheh, Raymond K., et al.
Published: (2025) -
Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
by: Patir, Rupam, et al.
Published: (2025)