DarkBench: Benchmarking Dark Patterns in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866929759356715008 |
|---|---|
| author | Kran, Esben Nguyen, Hieu Minh "Jord" Kundu, Akash Jawhar, Sami Park, Jinsuk Jurewicz, Mateusz Maria |
| author_facet | Kran, Esben Nguyen, Hieu Minh "Jord" Kundu, Akash Jawhar, Sami Park, Jinsuk Jurewicz, Mateusz Maria |
| contents | We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs). Our benchmark comprises 660 prompts across six categories: brand bias, user retention, sycophancy, anthropomorphism, harmful generation, and sneaking. We evaluate models from five leading companies (OpenAI, Anthropic, Meta, Mistral, Google) and find that some LLMs are explicitly designed to favor their developers' products and exhibit untruthful communication, among other manipulative behaviors. Companies developing LLMs should recognize and mitigate the impact of dark design patterns to promote more ethical AI. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_10728 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DarkBench: Benchmarking Dark Patterns in Large Language Models Kran, Esben Nguyen, Hieu Minh "Jord" Kundu, Akash Jawhar, Sami Park, Jinsuk Jurewicz, Mateusz Maria Computation and Language Artificial Intelligence Computers and Society We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs). Our benchmark comprises 660 prompts across six categories: brand bias, user retention, sycophancy, anthropomorphism, harmful generation, and sneaking. We evaluate models from five leading companies (OpenAI, Anthropic, Meta, Mistral, Google) and find that some LLMs are explicitly designed to favor their developers' products and exhibit untruthful communication, among other manipulative behaviors. Companies developing LLMs should recognize and mitigate the impact of dark design patterns to promote more ethical AI. |
| title | DarkBench: Benchmarking Dark Patterns in Large Language Models |
| topic | Computation and Language Artificial Intelligence Computers and Society |
| url | https://arxiv.org/abs/2503.10728 |