Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Zeren, Peng, Zifan, Liu, Yule, Sun, Zhen, Li, Mingchen, Zheng, Jingyi, He, Xinlei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
by: Peng, Zifan, et al.
Published: (2025)
by: Peng, Zifan, et al.
Published: (2025)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
by: Liu, Yule, et al.
Published: (2026)
by: Liu, Yule, et al.
Published: (2026)
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
"What Did It Actually Do?": Understanding Risk Awareness and Traceability for Computer-Use Agents
by: Peng, Zifan, et al.
Published: (2026)
by: Peng, Zifan, et al.
Published: (2026)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
by: Liu, Yule, et al.
Published: (2025)
by: Liu, Yule, et al.
Published: (2025)
Propagating Unsafe Actions in LLM Controlled Multi-Robot Collaboration via Single Robot Compromise
by: Huang, Zhen, et al.
Published: (2026)
by: Huang, Zhen, et al.
Published: (2026)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
Distortion Search, A Web Search Privacy Heuristic
by: Mivule, Kato, et al.
Published: (2025)
by: Mivule, Kato, et al.
Published: (2025)
DeepTx: Real-Time Transaction Risk Analysis via Multi-Modal Features and LLM Reasoning
by: Liu, Yixuan, et al.
Published: (2025)
by: Liu, Yixuan, et al.
Published: (2025)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Fast Summary-based Whole-program Analysis to Identify Unsafe Memory Accesses in Rust
by: Zhou, Jie, et al.
Published: (2023)
by: Zhou, Jie, et al.
Published: (2023)
RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing
by: Zhang, Wenhui, et al.
Published: (2026)
by: Zhang, Wenhui, et al.
Published: (2026)
Vital: Vulnerability-Oriented Symbolic Execution via Type-Unsafe Pointer-Guided Monte Carlo Tree Search
by: Tu, Haoxin, et al.
Published: (2024)
by: Tu, Haoxin, et al.
Published: (2024)
Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models
by: Luo, Weidi, et al.
Published: (2026)
by: Luo, Weidi, et al.
Published: (2026)
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
Friend or Foe Inside? Exploring In-Process Isolation to Maintain Memory Safety for Unsafe Rust
by: Gülmez, Merve, et al.
Published: (2023)
by: Gülmez, Merve, et al.
Published: (2023)
Searching for Privacy Risks in LLM Agents via Simulation
by: Zhang, Yanzhe, et al.
Published: (2025)
by: Zhang, Yanzhe, et al.
Published: (2025)
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024)
by: Goyal, Shubh, et al.
Published: (2024)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses
by: Sun, Zhen, et al.
Published: (2026)
by: Sun, Zhen, et al.
Published: (2026)
Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
Targeted Fuzzing for Unsafe Rust Code: Leveraging Selective Instrumentation
by: Paaßen, David, et al.
Published: (2025)
by: Paaßen, David, et al.
Published: (2025)
Towards Principled Analysis and Mitigation of Space Cyber Risks
by: Ear, Ekzhin
Published: (2025)
by: Ear, Ekzhin
Published: (2025)
RiskTagger: An LLM-based Agent for Automatic Annotation of Web3 Crypto Money Laundering Behaviors
by: Lin, Dan, et al.
Published: (2025)
by: Lin, Dan, et al.
Published: (2025)
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
by: He, Sinan, et al.
Published: (2025)
by: He, Sinan, et al.
Published: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Quantitative Analysis of UAV Intrusion Mitigation for Border Security in 5G with LEO Backhaul Impairments
by: Upadhyay, Rajendra, et al.
Published: (2025)
by: Upadhyay, Rajendra, et al.
Published: (2025)
NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
by: Asl, Javad Rafiei, et al.
Published: (2025)
by: Asl, Javad Rafiei, et al.
Published: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Rethinking Side-Channel Analysis: Automated Discovery and Analysis of Side-Channel Leakage with LLM-Assisted Agents
by: Xu, Zhen, et al.
Published: (2026)
by: Xu, Zhen, et al.
Published: (2026)
Privacy-Aware White and Black List Searching for Fraud Analysis
by: Buchanan, William J, et al.
Published: (2025)
by: Buchanan, William J, et al.
Published: (2025)
Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation
by: Andreucci, Biagio, et al.
Published: (2026)
by: Andreucci, Biagio, et al.
Published: (2026)
Investigating the Impact of Dark Patterns on LLM-Based Web Agents
by: Ersoy, Devin, et al.
Published: (2025)
by: Ersoy, Devin, et al.
Published: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
by: Li, MingSheng, et al.
Published: (2025)
by: Li, MingSheng, et al.
Published: (2025)
Similar Items
-
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
by: Peng, Zifan, et al.
Published: (2025) -
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025) -
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
by: Liu, Yule, et al.
Published: (2026) -
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024) -
"What Did It Actually Do?": Understanding Risk Awareness and Traceability for Computer-Use Agents
by: Peng, Zifan, et al.
Published: (2026)