Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Shawn Wanxiang, Liao, Junxuan, Liu, Jing, Zheng, Mai, Arpaci-Dusseau, Andrea C., Arpaci-Dusseau, Remzi H.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914478819377152
author Zhong, Shawn Wanxiang
Liao, Junxuan
Liu, Jing
Zheng, Mai
Arpaci-Dusseau, Andrea C.
Arpaci-Dusseau, Remzi H.
author_facet Zhong, Shawn Wanxiang
Liao, Junxuan
Liu, Jing
Zheng, Mai
Arpaci-Dusseau, Andrea C.
Arpaci-Dusseau, Remzi H.
contents AI coding agents operate directly on users' filesystems, where they regularly corrupt data, delete files, and leak secrets. Current approaches force a tradeoff between safety and autonomy: unrestricted access risks harm, while frequent permission prompts burden users and block agents. To understand this problem, we conduct the first systematic study of agent filesystem misuse, analyzing 290 public reports across 13 frameworks. Our analysis reveals that today's agents have limited information about their filesystem effects and insufficient control over them. We therefore argue for shifting this information and control to the filesystem itself. Based on this principle, we design YoloFS, an agent-native filesystem with three techniques. Staging isolates all mutations before commit, giving users corrective control. Snapshots extend this control to agents, letting them detect and correct their own mistakes. Progressive permission provides users with preventive control by gating access with minimal interaction. To evaluate YoloFS, we introduce a new methodology that captures user-agent-filesystem interactions. On 11 tasks with hidden side effects, YoloFS enables agent self-correction in 8 while keeping all effects staged and reviewable. On 112 routine tasks, YoloFS requires fewer user interactions while matching the baseline success rate.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13536
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy
Zhong, Shawn Wanxiang
Liao, Junxuan
Liu, Jing
Zheng, Mai
Arpaci-Dusseau, Andrea C.
Arpaci-Dusseau, Remzi H.
Operating Systems
AI coding agents operate directly on users' filesystems, where they regularly corrupt data, delete files, and leak secrets. Current approaches force a tradeoff between safety and autonomy: unrestricted access risks harm, while frequent permission prompts burden users and block agents. To understand this problem, we conduct the first systematic study of agent filesystem misuse, analyzing 290 public reports across 13 frameworks. Our analysis reveals that today's agents have limited information about their filesystem effects and insufficient control over them. We therefore argue for shifting this information and control to the filesystem itself. Based on this principle, we design YoloFS, an agent-native filesystem with three techniques. Staging isolates all mutations before commit, giving users corrective control. Snapshots extend this control to agents, letting them detect and correct their own mistakes. Progressive permission provides users with preventive control by gating access with minimal interaction. To evaluate YoloFS, we introduce a new methodology that captures user-agent-filesystem interactions. On 11 tasks with hidden side effects, YoloFS enables agent self-correction in 8 while keeping all effects staged and reviewable. On 112 routine tasks, YoloFS requires fewer user interactions while matching the baseline success rate.
title Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy
topic Operating Systems
url https://arxiv.org/abs/2604.13536