SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Jianshuo, Guo, Sheng, Wang, Hao, Chen, Xun, Liu, Zhuotao, Zhang, Tianwei, Xu, Ke, Huang, Minlie, Qiu, Han
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918529261895680
author Dong, Jianshuo
Guo, Sheng
Wang, Hao
Chen, Xun
Liu, Zhuotao
Zhang, Tianwei
Xu, Ke
Huang, Minlie
Qiu, Han
author_facet Dong, Jianshuo
Guo, Sheng
Wang, Hao
Chen, Xun
Liu, Zhuotao
Zhang, Tianwei
Xu, Ke
Huang, Minlie
Qiu, Han
contents Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new threat surface: unreliable search results can mislead agents into producing unsafe outputs. Real-world incidents and our two in-the-wild observations show that such failures can occur in practice. To study this threat systematically, we propose SafeSearch, an automated red-teaming framework that is scalable, cost-efficient, and lightweight, enabling sandboxed safety evaluation of search agents. Using this, we generate 300 test cases spanning five risk categories (e.g., misinformation and prompt injection) and evaluate three search agent scaffolds across 17 representative LLMs. Our results reveal substantial vulnerabilities in LLM-based search agents, with the highest ASR reaching 90.5% for GPT-4.1-mini in a search-workflow setting. Moreover, we find that common defenses, such as reminder prompting, offer limited protection. Overall, SafeSearch provides a practical way to measure and improve the safety of LLM-based search agents.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
Dong, Jianshuo
Guo, Sheng
Wang, Hao
Chen, Xun
Liu, Zhuotao
Zhang, Tianwei
Xu, Ke
Huang, Minlie
Qiu, Han
Artificial Intelligence
Computation and Language
Cryptography and Security
Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new threat surface: unreliable search results can mislead agents into producing unsafe outputs. Real-world incidents and our two in-the-wild observations show that such failures can occur in practice. To study this threat systematically, we propose SafeSearch, an automated red-teaming framework that is scalable, cost-efficient, and lightweight, enabling sandboxed safety evaluation of search agents. Using this, we generate 300 test cases spanning five risk categories (e.g., misinformation and prompt injection) and evaluate three search agent scaffolds across 17 representative LLMs. Our results reveal substantial vulnerabilities in LLM-based search agents, with the highest ASR reaching 90.5% for GPT-4.1-mini in a search-workflow setting. Moreover, we find that common defenses, such as reminder prompting, offer limited protection. Overall, SafeSearch provides a practical way to measure and improve the safety of LLM-based search agents.
title SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
topic Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2509.23694