Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Narad, Reuben, Suresh, Siddharth, Chen, Jiayi, Dysart-Bricken, Pine S. L., Mankoff, Bob, Nowak, Robert, Zhang, Jifan, Jain, Lalit
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908470802907136
author Narad, Reuben
Suresh, Siddharth
Chen, Jiayi
Dysart-Bricken, Pine S. L.
Mankoff, Bob
Nowak, Robert
Zhang, Jifan
Jain, Lalit
author_facet Narad, Reuben
Suresh, Siddharth
Chen, Jiayi
Dysart-Bricken, Pine S. L.
Mankoff, Bob
Nowak, Robert
Zhang, Jifan
Jain, Lalit
contents We present HumorBench, a benchmark designed to evaluate large language models' (LLMs) ability to reason about and explain sophisticated humor in cartoon captions. As reasoning models increasingly saturate existing benchmarks in mathematics and science, novel and challenging evaluations of model intelligence beyond STEM domains are essential. Reasoning is fundamentally involved in text-based humor comprehension, requiring the identification of connections between concepts in cartoons/captions and external cultural references, wordplays, and other mechanisms. HumorBench includes approximately 300 unique cartoon-caption pairs from the New Yorker Caption Contest and Cartoonstock.com, with expert-annotated evaluation rubrics identifying essential joke elements. LLMs are evaluated based on their explanations towards the humor and abilities in identifying the joke elements. To perform well on this task, models must form and test hypotheses about associations between concepts, potentially backtracking from initial interpretations to arrive at the most plausible explanation. Our extensive benchmarking of current SOTA models reveals three key insights: (1) LLM progress on STEM reasoning transfers effectively to humor comprehension; (2) models trained exclusively on STEM reasoning data still perform well on HumorBench, demonstrating strong transferability of reasoning abilities; and (3) test-time scaling by increasing thinking token budgets yields mixed results across different models in humor reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
Narad, Reuben
Suresh, Siddharth
Chen, Jiayi
Dysart-Bricken, Pine S. L.
Mankoff, Bob
Nowak, Robert
Zhang, Jifan
Jain, Lalit
Computation and Language
Artificial Intelligence
We present HumorBench, a benchmark designed to evaluate large language models' (LLMs) ability to reason about and explain sophisticated humor in cartoon captions. As reasoning models increasingly saturate existing benchmarks in mathematics and science, novel and challenging evaluations of model intelligence beyond STEM domains are essential. Reasoning is fundamentally involved in text-based humor comprehension, requiring the identification of connections between concepts in cartoons/captions and external cultural references, wordplays, and other mechanisms. HumorBench includes approximately 300 unique cartoon-caption pairs from the New Yorker Caption Contest and Cartoonstock.com, with expert-annotated evaluation rubrics identifying essential joke elements. LLMs are evaluated based on their explanations towards the humor and abilities in identifying the joke elements. To perform well on this task, models must form and test hypotheses about associations between concepts, potentially backtracking from initial interpretations to arrive at the most plausible explanation. Our extensive benchmarking of current SOTA models reveals three key insights: (1) LLM progress on STEM reasoning transfers effectively to humor comprehension; (2) models trained exclusively on STEM reasoning data still perform well on HumorBench, demonstrating strong transferability of reasoning abilities; and (3) test-time scaling by increasing thinking token budgets yields mixed results across different models in humor reasoning.
title Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.21476