Tracking Capabilities for Safer Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Odersky, Martin, Zhao, Yaoyu, Xu, Yichen, Bračevac, Oliver, Pham, Cao Nguyen
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917541624938496
author Odersky, Martin
Zhao, Yaoyu
Xu, Yichen
Bračevac, Oliver
Pham, Cao Nguyen
author_facet Odersky, Martin
Zhao, Yaoyu
Xu, Yichen
Bračevac, Oliver
Pham, Cao Nguyen
contents AI agents that interact with the real world through tool calls pose fundamental safety challenges: agents might leak private information, cause unintended side effects, or be manipulated through prompt injection. To address these challenges, we propose to put the agent in a programming-language-based "safety harness": instead of calling tools directly, agents express their intentions as code in a capability-safe language: Scala 3 with capture checking. Capabilities are program variables that regulate access to effects and resources of interest. Scala's type system tracks capabilities statically, providing fine-grained control over what an agent can do. In particular, it enables local purity, the ability to enforce that sub-computations are side-effect-free, preventing information leakage when agents process classified data. We demonstrate that extensible agent safety harnesses can be built by leveraging a strong type system with tracked capabilities. Our experiments show that agents can generate capability-safe code with no significant loss in task performance, while the type system reliably prevents unsafe behaviors such as information leakage and malicious side effects.
format Preprint
id arxiv_https___arxiv_org_abs_2603_00991
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tracking Capabilities for Safer Agents
Odersky, Martin
Zhao, Yaoyu
Xu, Yichen
Bračevac, Oliver
Pham, Cao Nguyen
Artificial Intelligence
Programming Languages
AI agents that interact with the real world through tool calls pose fundamental safety challenges: agents might leak private information, cause unintended side effects, or be manipulated through prompt injection. To address these challenges, we propose to put the agent in a programming-language-based "safety harness": instead of calling tools directly, agents express their intentions as code in a capability-safe language: Scala 3 with capture checking. Capabilities are program variables that regulate access to effects and resources of interest. Scala's type system tracks capabilities statically, providing fine-grained control over what an agent can do. In particular, it enables local purity, the ability to enforce that sub-computations are side-effect-free, preventing information leakage when agents process classified data. We demonstrate that extensible agent safety harnesses can be built by leveraging a strong type system with tracked capabilities. Our experiments show that agents can generate capability-safe code with no significant loss in task performance, while the type system reliably prevents unsafe behaviors such as information leakage and malicious side effects.
title Tracking Capabilities for Safer Agents
topic Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2603.00991