Exploring the Secondary Risks of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiawei, Fang, Zhengwei, Tian, Yu, Du, Jiawei, Yu, Chao, Yin, Zhaoxia, Su, Hang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911624256815104
author Chen, Jiawei
Fang, Zhengwei
Tian, Yu
Du, Jiawei
Yu, Chao
Yin, Zhaoxia
Su, Hang
author_facet Chen, Jiawei
Fang, Zhengwei
Tian, Yu
Du, Jiawei
Yu, Chao
Yin, Zhaoxia
Su, Hang
contents Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior research has primarily focused on jailbreak attacks, less attention has been given to non-adversarial failures that subtly emerge during benign interactions. We introduce secondary risks a novel class of failure modes marked by harmful or misleading behaviors during benign prompts. Unlike adversarial attacks, these risks stem from imperfect generalization and often evade standard safety mechanisms. To enable systematic evaluation, we introduce two risk primitives verbose response and speculative advice that capture the core failure patterns. Building on these definitions, we propose SecLens, a black-box, multi-objective search framework that efficiently elicits secondary risk behaviors by optimizing task relevance, risk activation, and linguistic plausibility. To support reproducible evaluation, we release SecRiskBench, a benchmark dataset of 650 prompts covering eight diverse real-world risk categories. Experimental results from extensive evaluations on 16 popular models demonstrate that secondary risks are widespread, transferable across models, and modality independent, emphasizing the urgent need for enhanced safety mechanisms to address benign yet harmful LLM behaviors in real-world deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12382
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Secondary Risks of Large Language Models
Chen, Jiawei
Fang, Zhengwei
Tian, Yu
Du, Jiawei
Yu, Chao
Yin, Zhaoxia
Su, Hang
Machine Learning
Artificial Intelligence
Cryptography and Security
Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior research has primarily focused on jailbreak attacks, less attention has been given to non-adversarial failures that subtly emerge during benign interactions. We introduce secondary risks a novel class of failure modes marked by harmful or misleading behaviors during benign prompts. Unlike adversarial attacks, these risks stem from imperfect generalization and often evade standard safety mechanisms. To enable systematic evaluation, we introduce two risk primitives verbose response and speculative advice that capture the core failure patterns. Building on these definitions, we propose SecLens, a black-box, multi-objective search framework that efficiently elicits secondary risk behaviors by optimizing task relevance, risk activation, and linguistic plausibility. To support reproducible evaluation, we release SecRiskBench, a benchmark dataset of 650 prompts covering eight diverse real-world risk categories. Experimental results from extensive evaluations on 16 popular models demonstrate that secondary risks are widespread, transferable across models, and modality independent, emphasizing the urgent need for enhanced safety mechanisms to address benign yet harmful LLM behaviors in real-world deployments.
title Exploring the Secondary Risks of Large Language Models
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2506.12382