Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
Fuente:
arXiv
Saved in:
| Main Authors: | Narad, Reuben, Suresh, Siddharth, Chen, Jiayi, Dysart-Bricken, Pine S. L., Mankoff, Bob, Nowak, Robert, Zhang, Jifan, Jain, Lalit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs
by: Zhou, Kuan Lok, et al.
Published: (2025)
by: Zhou, Kuan Lok, et al.
Published: (2025)
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
by: Zhang, Jifan, et al.
Published: (2024)
by: Zhang, Jifan, et al.
Published: (2024)
Improved Algorithm for Deep Active Learning under Imbalance via Optimal Separation
by: Nuggehalli, Shyam, et al.
Published: (2023)
by: Nuggehalli, Shyam, et al.
Published: (2023)
Probing Neural TSP Representations for Prescriptive Decision Support
by: Narad, Reuben, et al.
Published: (2026)
by: Narad, Reuben, et al.
Published: (2026)
Learning to Actively Learn: A Robust Approach
by: Zhang, Jifan, et al.
Published: (2020)
by: Zhang, Jifan, et al.
Published: (2020)
One Joke to Rule them All? On the (Im)possibility of Generalizing Humor
by: Turgeman, Mor, et al.
Published: (2025)
by: Turgeman, Mor, et al.
Published: (2025)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Mechanistic Interpretability for Neural TSP Solvers
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
by: Vural, Hatice Merve, et al.
Published: (2026)
by: Vural, Hatice Merve, et al.
Published: (2026)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
Humor in the Academic Library: You Must Be Joking! or, How Many Academic Librarians Does It Take To Change a Lightbulb?
by: Black, Leah, et al.
Published: (1999)
by: Black, Leah, et al.
Published: (1999)
Physicists Are Still Joking
by: Halperin, Igor
Published: (2025)
by: Halperin, Igor
Published: (2025)
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
by: Arabelly, Abhinav, et al.
Published: (2025)
by: Arabelly, Abhinav, et al.
Published: (2025)
The Language of Jokes in the Digital Age
by: Chiaro, Delia
Published: (2025)
by: Chiaro, Delia
Published: (2025)
My Favorite Math Jokes
by: Khovanova, Tanya
Published: (2024)
by: Khovanova, Tanya
Published: (2024)
Gag Gifting: The Joke and the Poke
by: Robert M. Schindler, et al.
Published: (2025)
by: Robert M. Schindler, et al.
Published: (2025)
Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models
by: Horvitz, Zachary, et al.
Published: (2024)
by: Horvitz, Zachary, et al.
Published: (2024)
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs
by: Wang, Zhizhi, et al.
Published: (2026)
by: Wang, Zhizhi, et al.
Published: (2026)
AHA: Human-Assisted Out-of-Distribution Generalization and Detection
by: Bai, Haoyue, et al.
Published: (2024)
by: Bai, Haoyue, et al.
Published: (2024)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
by: Tripathi, Sahil, et al.
Published: (2026)
by: Tripathi, Sahil, et al.
Published: (2026)
Business Ethics and Critical Consultant Jokes
by: Bouwmeester, Onno
Published: (2022)
by: Bouwmeester, Onno
Published: (2022)
The Middle East and the Ukraine War: Between Fear and Opportunity
by: Jeffrey Mankoff
Published: (2024)
by: Jeffrey Mankoff
Published: (2024)
RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
by: He, Xuming, et al.
Published: (2025)
by: He, Xuming, et al.
Published: (2025)
Humor Mechanics: Advancing Humor Generation with Multistep Reasoning
by: Tikhonov, Alexey, et al.
Published: (2024)
by: Tikhonov, Alexey, et al.
Published: (2024)
Assessing Innovation in Corporate and Government Libraries
by: Zeeman, Deane, et al.
Published: (2011)
by: Zeeman, Deane, et al.
Published: (2011)
Tools for the Future: Recreating or "Renovating" Information Services Using New Technologies.
by: Dysart, Jane I, et al.
Published: (1995)
by: Dysart, Jane I, et al.
Published: (1995)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
A remark on a theorem of Narasimhan and Ramanan
by: Pine, Jagadish
Published: (2024)
by: Pine, Jagadish
Published: (2024)
“Tu eres gallo... pero la de los huevos soy yo”: Producción y género en las maquiladoras de Honduras
by: Adrienne Pine
Published: (2009)
by: Adrienne Pine
Published: (2009)
A Basic Collection of Software for Children.
by: Pine, Susan
Published: (1991)
by: Pine, Susan
Published: (1991)
Tegucigolpe: donde se cruzan los caminos, se unen fronteras y divergen las percepciones
by: Adrienne Pine
Published: (2011)
by: Adrienne Pine
Published: (2011)
Low complexity binary words avoiding $(5/2)^+$-powers
by: Currie, James, et al.
Published: (2025)
by: Currie, James, et al.
Published: (2025)
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)
by: Naik, Ranjita, et al.
Published: (2023)
Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications
by: Kaushik, Saurabh, et al.
Published: (2026)
by: Kaushik, Saurabh, et al.
Published: (2026)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
by: Fernandez, Nigel, et al.
Published: (2025)
by: Fernandez, Nigel, et al.
Published: (2025)
Turing Jest: Distributional Semantics and One‐Line Jokes
by: Sean Trott, et al.
Published: (2025)
by: Sean Trott, et al.
Published: (2025)
Uncovering the Computational Ingredients of Human-Like Representations in LLMs
by: Studdiford, Zach, et al.
Published: (2025)
by: Studdiford, Zach, et al.
Published: (2025)
Deep Active Learning in the Open World
by: Xie, Tian, et al.
Published: (2024)
by: Xie, Tian, et al.
Published: (2024)
Similar Items
-
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs
by: Zhou, Kuan Lok, et al.
Published: (2025) -
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
by: Zhang, Jifan, et al.
Published: (2024) -
Improved Algorithm for Deep Active Learning under Imbalance via Optimal Separation
by: Nuggehalli, Shyam, et al.
Published: (2023) -
Probing Neural TSP Representations for Prescriptive Decision Support
by: Narad, Reuben, et al.
Published: (2026) -
Learning to Actively Learn: A Robust Approach
by: Zhang, Jifan, et al.
Published: (2020)