Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fettach, Yousra, Bied, Guillaume, Toivonen, Hannu, De Bie, Tijl
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911580502884352
author Fettach, Yousra
Bied, Guillaume
Toivonen, Hannu
De Bie, Tijl
author_facet Fettach, Yousra
Bied, Guillaume
Toivonen, Hannu
De Bie, Tijl
contents Humor is one of the most culturally embedded and socially significant dimensions of human communication, yet it remains largely unexplored as a dimension of Large Language Model (LLM) alignment. In this study, five frontier language models play the same Cards Against Humanity games (CAH) as human players. The models select the funniest response from a slate of ten candidate cards across 9,894 rounds. While all models exceed the random baseline, alignment with human preference remains modest. More striking is that models agree with each other substantially more often than they agree with humans. We show that this preference is partly explained by systematic position biases and content preferences, raising the question whether LLM humor judgment reflects genuine preference or structural artifacts of inference and alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08757
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
Fettach, Yousra
Bied, Guillaume
Toivonen, Hannu
De Bie, Tijl
Computation and Language
Artificial Intelligence
Humor is one of the most culturally embedded and socially significant dimensions of human communication, yet it remains largely unexplored as a dimension of Large Language Model (LLM) alignment. In this study, five frontier language models play the same Cards Against Humanity games (CAH) as human players. The models select the funniest response from a slate of ten candidate cards across 9,894 rounds. While all models exceed the random baseline, alignment with human preference remains modest. More striking is that models agree with each other substantially more often than they agree with humans. We show that this preference is partly explained by systematic position biases and content preferences, raising the question whether LLM humor judgment reflects genuine preference or structural artifacts of inference and alignment.
title Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.08757