DRAFT: Task Decoupled Latent Reasoning for Agent Safety

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Lin, Fang, Junfeng, Zhang, Dan, Shen, Fei, Wang, Xiang, Chua, Tat-Seng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913003965775872
author Wang, Lin
Fang, Junfeng
Zhang, Dan
Shen, Fei
Wang, Xiang
Chua, Tat-Seng
author_facet Wang, Lin
Fang, Junfeng
Zhang, Dan
Shen, Fei
Wang, Xiang
Chua, Tat-Seng
contents The advent of tool-using LLM agents shifts safety monitoring from output moderation to auditing long, noisy interaction trajectories, where risk-critical evidence is sparse-making standard binary supervision poorly suited for credit assignment. To address this, we propose DRAFT (Task Decoupled Latent Reasoning for Agent Safety), a latent reasoning framework that decouples safety judgment into two trainable stages: an Extractor that distills the full trajectory into a compact continuous latent draft, and a Reasoner that jointly attends to the draft and the original trajectory to predict safety. DRAFT avoids lossy explicit summarize-then-judge pipelines by performing evidence aggregation in latent space, enabling end-to-end differentiable training.Across benchmarks including ASSEBench and R-Judge, DRAFT consistently outperforms strong baselines, improving accuracy from 63.27% (LoRA) to 91.18% averaged over benchmarks, and learns more separable representations. Ablations demonstrate a clear synergy between the Extractor and the Reasoner.Overall, DRAFT suggests that continuous latent reasoning prior to readout is a practical path to robust agent safety under long-context supervision with sparse evidence.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03242
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DRAFT: Task Decoupled Latent Reasoning for Agent Safety
Wang, Lin
Fang, Junfeng
Zhang, Dan
Shen, Fei
Wang, Xiang
Chua, Tat-Seng
Machine Learning
The advent of tool-using LLM agents shifts safety monitoring from output moderation to auditing long, noisy interaction trajectories, where risk-critical evidence is sparse-making standard binary supervision poorly suited for credit assignment. To address this, we propose DRAFT (Task Decoupled Latent Reasoning for Agent Safety), a latent reasoning framework that decouples safety judgment into two trainable stages: an Extractor that distills the full trajectory into a compact continuous latent draft, and a Reasoner that jointly attends to the draft and the original trajectory to predict safety. DRAFT avoids lossy explicit summarize-then-judge pipelines by performing evidence aggregation in latent space, enabling end-to-end differentiable training.Across benchmarks including ASSEBench and R-Judge, DRAFT consistently outperforms strong baselines, improving accuracy from 63.27% (LoRA) to 91.18% averaged over benchmarks, and learns more separable representations. Ablations demonstrate a clear synergy between the Extractor and the Reasoner.Overall, DRAFT suggests that continuous latent reasoning prior to readout is a practical path to robust agent safety under long-context supervision with sparse evidence.
title DRAFT: Task Decoupled Latent Reasoning for Agent Safety
topic Machine Learning
url https://arxiv.org/abs/2604.03242