COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Wenkai, Zhou, Pengyang, Xu, Jiahe, Qian, Jiaming, He, Haozhe, Huang, Zhihao, Chen, Chaochao, Zheng, Xiaolin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910271649349632
author Shen, Wenkai
Zhou, Pengyang
Xu, Jiahe
Qian, Jiaming
He, Haozhe
Huang, Zhihao
Chen, Chaochao
Zheng, Xiaolin
author_facet Shen, Wenkai
Zhou, Pengyang
Xu, Jiahe
Qian, Jiaming
He, Haozhe
Huang, Zhihao
Chen, Chaochao
Zheng, Xiaolin
contents LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmful intents may decompose into seemingly innocuous sub-queries that lead to unsafe outcomes. Existing alignment methods struggle to capture sparse safety signals and fail to supervise diverse violations across multi-step interactions. We propose COMPASS, a Cognitive MCTS-Guided Process Alignment framework designed to achieve robust safety alignment throughout the agent workflow while preserving general utility. COMPASS integrates cognitive tree exploration (CTE) to efficiently synthesize stealthy attack trajectories, and introspective step-wise alignment (ISA) to isolate risky intermediate actions for fine-grained process supervision. Empirical results show that COMPASS achieves a favorable safety-utility trade-off while requiring substantially less training data.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30838
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents
Shen, Wenkai
Zhou, Pengyang
Xu, Jiahe
Qian, Jiaming
He, Haozhe
Huang, Zhihao
Chen, Chaochao
Zheng, Xiaolin
Artificial Intelligence
LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmful intents may decompose into seemingly innocuous sub-queries that lead to unsafe outcomes. Existing alignment methods struggle to capture sparse safety signals and fail to supervise diverse violations across multi-step interactions. We propose COMPASS, a Cognitive MCTS-Guided Process Alignment framework designed to achieve robust safety alignment throughout the agent workflow while preserving general utility. COMPASS integrates cognitive tree exploration (CTE) to efficiently synthesize stealthy attack trajectories, and introspective step-wise alignment (ISA) to isolate risky intermediate actions for fine-grained process supervision. Empirical results show that COMPASS achieves a favorable safety-utility trade-off while requiring substantially less training data.
title COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2605.30838