Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Atil, Berk, Gupta, Vipul, Das, Sarkar Snigdha Sarathi, Passonneau, Rebecca J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913800724152320
author Atil, Berk
Gupta, Vipul
Das, Sarkar Snigdha Sarathi
Passonneau, Rebecca J.
author_facet Atil, Berk
Gupta, Vipul
Das, Sarkar Snigdha Sarathi
Passonneau, Rebecca J.
contents Large language models (LLMs) have become ubiquitous, thus it is important to understand their risks and limitations. Smaller LLMs can be deployed where compute resources are constrained, such as edge devices, but with different propensity to generate harmful output. Mitigation of LLM harm typically depends on annotating the harmfulness of LLM output, which is expensive to collect from humans. This work studies two questions: How do smaller LLMs rank regarding generation of harmful content? How well can larger LLMs annotate harmfulness? We prompt three small LLMs to elicit harmful content of various types, such as discriminatory language, offensive content, privacy invasion, or negative influence, and collect human rankings of their outputs. Then, we evaluate three state-of-the-art large LLMs on their ability to annotate the harmfulness of these responses. We find that the smaller models differ with respect to harmfulness. We also find that large LLMs show low to moderate agreement with humans. These findings underline the need for further work on harm mitigation in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
Atil, Berk
Gupta, Vipul
Das, Sarkar Snigdha Sarathi
Passonneau, Rebecca J.
Computation and Language
Large language models (LLMs) have become ubiquitous, thus it is important to understand their risks and limitations. Smaller LLMs can be deployed where compute resources are constrained, such as edge devices, but with different propensity to generate harmful output. Mitigation of LLM harm typically depends on annotating the harmfulness of LLM output, which is expensive to collect from humans. This work studies two questions: How do smaller LLMs rank regarding generation of harmful content? How well can larger LLMs annotate harmfulness? We prompt three small LLMs to elicit harmful content of various types, such as discriminatory language, offensive content, privacy invasion, or negative influence, and collect human rankings of their outputs. Then, we evaluate three state-of-the-art large LLMs on their ability to annotate the harmfulness of these responses. We find that the smaller models differ with respect to harmfulness. We also find that large LLMs show low to moderate agreement with humans. These findings underline the need for further work on harm mitigation in LLMs.
title Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
topic Computation and Language
url https://arxiv.org/abs/2502.05291