Mimicking How Humans Interpret Out-of-Context Sentences Through Controlled Toxicity Decoding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Trusca, Maria Mihaela, Allein, Liesbeth
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910133495267328
author Trusca, Maria Mihaela
Allein, Liesbeth
author_facet Trusca, Maria Mihaela
Allein, Liesbeth
contents Interpretations of a single sentence can vary, particularly when its context is lost. This paper aims to simulate how readers perceive content with varying toxicity levels by generating diverse interpretations of out-of-context sentences. By modeling toxicity, we can anticipate misunderstandings and reveal hidden toxic meanings. Our proposed decoding strategy explicitly controls toxicity in the set of generated interpretations by (i) aligning interpretation toxicity with the input, (ii) relaxing toxicity constraints for more toxic input sentences, and (iii) promoting diversity in toxicity levels within the set of generated interpretations. Experimental results show that our method improves alignment with human-written interpretations in both syntax and semantics while reducing model prediction uncertainty.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mimicking How Humans Interpret Out-of-Context Sentences Through Controlled Toxicity Decoding
Trusca, Maria Mihaela
Allein, Liesbeth
Computation and Language
Interpretations of a single sentence can vary, particularly when its context is lost. This paper aims to simulate how readers perceive content with varying toxicity levels by generating diverse interpretations of out-of-context sentences. By modeling toxicity, we can anticipate misunderstandings and reveal hidden toxic meanings. Our proposed decoding strategy explicitly controls toxicity in the set of generated interpretations by (i) aligning interpretation toxicity with the input, (ii) relaxing toxicity constraints for more toxic input sentences, and (iii) promoting diversity in toxicity levels within the set of generated interpretations. Experimental results show that our method improves alignment with human-written interpretations in both syntax and semantics while reducing model prediction uncertainty.
title Mimicking How Humans Interpret Out-of-Context Sentences Through Controlled Toxicity Decoding
topic Computation and Language
url https://arxiv.org/abs/2503.08159