Dynamic Evaluation for Oversensitivity in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pu, Sophia Xiao, Cheng, Sitao, Wang, Xin Eric, Wang, William Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917031989739520
author Pu, Sophia Xiao
Cheng, Sitao
Wang, Xin Eric
Wang, William Yang
author_facet Pu, Sophia Xiao
Cheng, Sitao
Wang, Xin Eric
Wang, William Yang
contents Oversensitivity occurs when language models defensively reject prompts that are actually benign. This behavior not only disrupts user interactions but also obscures the boundary between harmful and harmless content. Existing benchmarks rely on static datasets that degrade overtime as models evolve, leading to data contamination and diminished evaluative power. To address this, we develop a framework that dynamically generates model-specific challenging datasets, capturing emerging defensive patterns and aligning with each model's unique behavior. Building on this approach, we construct OVERBENCH, a benchmark that aggregates these datasets across diverse LLM families, encompassing 450,000 samples from 25 models. OVERBENCH provides a dynamic and evolving perspective on oversensitivity, allowing for continuous monitoring of defensive triggers as models advance, highlighting vulnerabilities that static datasets overlook.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Evaluation for Oversensitivity in LLMs
Pu, Sophia Xiao
Cheng, Sitao
Wang, Xin Eric
Wang, William Yang
Computation and Language
Oversensitivity occurs when language models defensively reject prompts that are actually benign. This behavior not only disrupts user interactions but also obscures the boundary between harmful and harmless content. Existing benchmarks rely on static datasets that degrade overtime as models evolve, leading to data contamination and diminished evaluative power. To address this, we develop a framework that dynamically generates model-specific challenging datasets, capturing emerging defensive patterns and aligning with each model's unique behavior. Building on this approach, we construct OVERBENCH, a benchmark that aggregates these datasets across diverse LLM families, encompassing 450,000 samples from 25 models. OVERBENCH provides a dynamic and evolving perspective on oversensitivity, allowing for continuous monitoring of defensive triggers as models advance, highlighting vulnerabilities that static datasets overlook.
title Dynamic Evaluation for Oversensitivity in LLMs
topic Computation and Language
url https://arxiv.org/abs/2510.19005