Reflection-Driven Control for Trustworthy Code Agents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Bin, Quan, Jiazheng, Yu, Xingrui, Hu, Hansen, Yuhao, Tsang, Ivor
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912788930101248
author Wang, Bin
Quan, Jiazheng
Yu, Xingrui
Hu, Hansen
Yuhao
Tsang, Ivor
author_facet Wang, Bin
Quan, Jiazheng
Yu, Xingrui
Hu, Hansen
Yuhao
Tsang, Ivor
contents Contemporary large language model (LLM) agents are remarkably capable, but they still lack reliable safety controls and can produce unconstrained, unpredictable, and even actively harmful outputs. To address this, we introduce Reflection-Driven Control, a standardized and pluggable control module that can be seamlessly integrated into general agent architectures. Reflection-Driven Control elevates "self-reflection" from a post hoc patch into an explicit step in the agent's own reasoning process: during generation, the agent continuously runs an internal reflection loop that monitors and evaluates its own decision path. When potential risks are detected, the system retrieves relevant repair examples and secure coding guidelines from an evolving reflective memory, injecting these evidence-based constraints directly into subsequent reasoning steps. We instantiate Reflection-Driven Control in the setting of secure code generation and systematically evaluate it across eight classes of security-critical programming tasks. Empirical results show that Reflection-Driven Control substantially improves the security and policy compliance of generated code while largely preserving functional correctness, with minimal runtime and token overhead. Taken together, these findings indicate that Reflection-Driven Control is a practical path toward trustworthy AI coding agents: it enables designs that are simultaneously autonomous, safer by construction, and auditable.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reflection-Driven Control for Trustworthy Code Agents
Wang, Bin
Quan, Jiazheng
Yu, Xingrui
Hu, Hansen
Yuhao
Tsang, Ivor
Cryptography and Security
Artificial Intelligence
Software Engineering
Contemporary large language model (LLM) agents are remarkably capable, but they still lack reliable safety controls and can produce unconstrained, unpredictable, and even actively harmful outputs. To address this, we introduce Reflection-Driven Control, a standardized and pluggable control module that can be seamlessly integrated into general agent architectures. Reflection-Driven Control elevates "self-reflection" from a post hoc patch into an explicit step in the agent's own reasoning process: during generation, the agent continuously runs an internal reflection loop that monitors and evaluates its own decision path. When potential risks are detected, the system retrieves relevant repair examples and secure coding guidelines from an evolving reflective memory, injecting these evidence-based constraints directly into subsequent reasoning steps. We instantiate Reflection-Driven Control in the setting of secure code generation and systematically evaluate it across eight classes of security-critical programming tasks. Empirical results show that Reflection-Driven Control substantially improves the security and policy compliance of generated code while largely preserving functional correctness, with minimal runtime and token overhead. Taken together, these findings indicate that Reflection-Driven Control is a practical path toward trustworthy AI coding agents: it enables designs that are simultaneously autonomous, safer by construction, and auditable.
title Reflection-Driven Control for Trustworthy Code Agents
topic Cryptography and Security
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2512.21354