SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhan, Qiusi, Budiman-Chan, Angeline, Zayed, Abdelrahman, Guo, Xingzhi, Kang, Daniel, Kim, Joo-Kyung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
Do Biased Models Have Biased Thoughts?
von: Rajwal, Swati, et al.
Veröffentlicht: (2025)
von: Rajwal, Swati, et al.
Veröffentlicht: (2025)
The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
Removing RLHF Protections in GPT-4 via Fine-Tuning
von: Zhan, Qiusi, et al.
Veröffentlicht: (2023)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2023)
Exploring Safety-Utility Trade-Offs in Personalized Language Models
von: Vijjini, Anvesh Rao, et al.
Veröffentlicht: (2024)
von: Vijjini, Anvesh Rao, et al.
Veröffentlicht: (2024)
Where to Search: Measure the Prior-Structured Search Space of LLM Agents
von: Song, Zhuo-Yang
Veröffentlicht: (2025)
von: Song, Zhuo-Yang
Veröffentlicht: (2025)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
von: Su, Zhe, et al.
Veröffentlicht: (2024)
von: Su, Zhe, et al.
Veröffentlicht: (2024)
AgentSquare: Automatic LLM Agent Search in Modular Design Space
von: Shang, Yu, et al.
Veröffentlicht: (2024)
von: Shang, Yu, et al.
Veröffentlicht: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Hierarchies over Vector Space: Orienting Word and Graph Embeddings
von: Guo, Xingzhi, et al.
Veröffentlicht: (2022)
von: Guo, Xingzhi, et al.
Veröffentlicht: (2022)
EvolveSearch: An Iterative Self-Evolving Search Agent
von: Zhang, Dingchu, et al.
Veröffentlicht: (2025)
von: Zhang, Dingchu, et al.
Veröffentlicht: (2025)
WaterSearch: Exploring Seed Pooling for Improving the Quality-Detectability Trade-off in LLM Watermarking
von: Lin, Yukang, et al.
Veröffentlicht: (2025)
von: Lin, Yukang, et al.
Veröffentlicht: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
When Is Enough Not Enough? Illusory Completion in Search Agents
von: Ko, Dayoon, et al.
Veröffentlicht: (2026)
von: Ko, Dayoon, et al.
Veröffentlicht: (2026)
Generative Subgraph Retrieval for Knowledge Graph-Grounded Dialog Generation
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
von: Qian, Lingfei, et al.
Veröffentlicht: (2025)
von: Qian, Lingfei, et al.
Veröffentlicht: (2025)
ZeroSearch: Incentivize the Search Capability of LLMs without Searching
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
LLM Tree Search
von: Wilson, Dylan
Veröffentlicht: (2024)
von: Wilson, Dylan
Veröffentlicht: (2024)
Searching for Privacy Risks in LLM Agents via Simulation
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
SE-Search: Self-Evolving Search Agent via Memory and Dense Reward
von: Li, Jian, et al.
Veröffentlicht: (2026)
von: Li, Jian, et al.
Veröffentlicht: (2026)
LLM Agents Improve Semantic Code Search
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
von: Lee, Gisang, et al.
Veröffentlicht: (2024)
von: Lee, Gisang, et al.
Veröffentlicht: (2024)
SearchLLM: Detecting LLM Paraphrased Text by Measuring the Similarity with Regeneration of the Candidate Source via Search Engine
von: Nguyen-Son, Hoang-Quoc, et al.
Veröffentlicht: (2026)
von: Nguyen-Son, Hoang-Quoc, et al.
Veröffentlicht: (2026)
LiteSearch: Efficacious Tree Search for LLM
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
von: Zhao, Qingfei, et al.
Veröffentlicht: (2025)
von: Zhao, Qingfei, et al.
Veröffentlicht: (2025)
Safe-Embed: Unveiling the Safety-Critical Knowledge of Sentence Encoders
von: Kim, Jinseok, et al.
Veröffentlicht: (2024)
von: Kim, Jinseok, et al.
Veröffentlicht: (2024)
LLM Agents can Autonomously Hack Websites
von: Fang, Richard, et al.
Veröffentlicht: (2024)
von: Fang, Richard, et al.
Veröffentlicht: (2024)
Advancing LLM Safe Alignment with Safety Representation Ranking
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
von: Suh, Joseph, et al.
Veröffentlicht: (2026)
von: Suh, Joseph, et al.
Veröffentlicht: (2026)
Tree Search for Language Model Agents
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
von: Zhang, Yaocheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaocheng, et al.
Veröffentlicht: (2025)
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025) -
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024) -
Do Biased Models Have Biased Thoughts?
von: Rajwal, Swati, et al.
Veröffentlicht: (2025) -
The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents
von: Tang, Yihong, et al.
Veröffentlicht: (2025) -
Removing RLHF Protections in GPT-4 via Fine-Tuning
von: Zhan, Qiusi, et al.
Veröffentlicht: (2023)