DarkBench: Benchmarking Dark Patterns in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kran, Esben, Nguyen, Hieu Minh "Jord", Kundu, Akash, Jawhar, Sami, Park, Jinsuk, Jurewicz, Mateusz Maria
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929759356715008
author Kran, Esben
Nguyen, Hieu Minh "Jord"
Kundu, Akash
Jawhar, Sami
Park, Jinsuk
Jurewicz, Mateusz Maria
author_facet Kran, Esben
Nguyen, Hieu Minh "Jord"
Kundu, Akash
Jawhar, Sami
Park, Jinsuk
Jurewicz, Mateusz Maria
contents We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs). Our benchmark comprises 660 prompts across six categories: brand bias, user retention, sycophancy, anthropomorphism, harmful generation, and sneaking. We evaluate models from five leading companies (OpenAI, Anthropic, Meta, Mistral, Google) and find that some LLMs are explicitly designed to favor their developers' products and exhibit untruthful communication, among other manipulative behaviors. Companies developing LLMs should recognize and mitigate the impact of dark design patterns to promote more ethical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10728
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DarkBench: Benchmarking Dark Patterns in Large Language Models
Kran, Esben
Nguyen, Hieu Minh "Jord"
Kundu, Akash
Jawhar, Sami
Park, Jinsuk
Jurewicz, Mateusz Maria
Computation and Language
Artificial Intelligence
Computers and Society
We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs). Our benchmark comprises 660 prompts across six categories: brand bias, user retention, sycophancy, anthropomorphism, harmful generation, and sneaking. We evaluate models from five leading companies (OpenAI, Anthropic, Meta, Mistral, Google) and find that some LLMs are explicitly designed to favor their developers' products and exhibit untruthful communication, among other manipulative behaviors. Companies developing LLMs should recognize and mitigate the impact of dark design patterns to promote more ethical AI.
title DarkBench: Benchmarking Dark Patterns in Large Language Models
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2503.10728