Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Sakabe, Ritsu, Kim, Hwichan, Hirasawa, Tosho, Komachi, Mamoru |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pruning Multilingual Large Language Models for Multilingual Inference
by: Kim, Hwichan, et al.
Published: (2024)
by: Kim, Hwichan, et al.
Published: (2024)
Oogiri-Master: Benchmarking Humor Understanding via Oogiri
by: Murakami, Soichiro, et al.
Published: (2025)
by: Murakami, Soichiro, et al.
Published: (2025)
Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
by: Zhang, Huaying, et al.
Published: (2025)
by: Zhang, Huaying, et al.
Published: (2025)
Exploring the Effects of Alignment on Numerical Bias in Large Language Models
by: Sato, Ayako, et al.
Published: (2026)
by: Sato, Ayako, et al.
Published: (2026)
Assessing the Capability of LLMs in Solving POSCOMP Questions
by: Viegas, Cayo, et al.
Published: (2025)
by: Viegas, Cayo, et al.
Published: (2025)
A Single Linear Layer Yields Task-Adapted Low-Rank Matrices
by: Kim, Hwichan, et al.
Published: (2024)
by: Kim, Hwichan, et al.
Published: (2024)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
Large Language Models Are State-of-the-Art Evaluator for Grammatical Error Correction
by: Kobayashi, Masamune, et al.
Published: (2024)
by: Kobayashi, Masamune, et al.
Published: (2024)
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
by: Ruan, Kai, et al.
Published: (2024)
by: Ruan, Kai, et al.
Published: (2024)
HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation
by: Zhang, Jiajun, et al.
Published: (2025)
by: Zhang, Jiajun, et al.
Published: (2025)
An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
by: Vyas, Kaustubh, et al.
Published: (2025)
by: Vyas, Kaustubh, et al.
Published: (2025)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
by: Kwon, Deuksin, et al.
Published: (2024)
by: Kwon, Deuksin, et al.
Published: (2024)
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
by: Fettach, Yousra, et al.
Published: (2026)
by: Fettach, Yousra, et al.
Published: (2026)
Evaluation of Multilingual LLMs Personalized Text Generation Capabilities Targeting Groups and Social-Media Platforms
by: Macko, Dominik
Published: (2026)
by: Macko, Dominik
Published: (2026)
Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Models
by: Song, Sangmin, et al.
Published: (2025)
by: Song, Sangmin, et al.
Published: (2025)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
by: Sharshar, Ahmed, et al.
Published: (2026)
by: Sharshar, Ahmed, et al.
Published: (2026)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
by: Jang, Kyochul, et al.
Published: (2025)
by: Jang, Kyochul, et al.
Published: (2025)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
by: Fu, Weiping, et al.
Published: (2024)
by: Fu, Weiping, et al.
Published: (2024)
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
by: Ando, Kenichiro, et al.
Published: (2023)
by: Ando, Kenichiro, et al.
Published: (2023)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
by: Wang, Hanbin, et al.
Published: (2025)
by: Wang, Hanbin, et al.
Published: (2025)
Who Laughs with Whom? Disentangling Influential Factors in Humor Preferences across User Clusters and LLMs
by: Murakami, Soichiro, et al.
Published: (2026)
by: Murakami, Soichiro, et al.
Published: (2026)
Exploring Chinese Humor Generation: A Study on Two-Part Allegorical Sayings
by: Xu, Rongwu
Published: (2024)
by: Xu, Rongwu
Published: (2024)
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
by: Li, Miao, et al.
Published: (2024)
by: Li, Miao, et al.
Published: (2024)
Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
One Joke to Rule them All? On the (Im)possibility of Generalizing Humor
by: Turgeman, Mor, et al.
Published: (2025)
by: Turgeman, Mor, et al.
Published: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
by: Yoo, Yongmin, et al.
Published: (2025)
by: Yoo, Yongmin, et al.
Published: (2025)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
by: Zhang, Leizhen, et al.
Published: (2026)
by: Zhang, Leizhen, et al.
Published: (2026)
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
by: Koo, Hamin, et al.
Published: (2025)
by: Koo, Hamin, et al.
Published: (2025)
Revisiting Meta-evaluation for Grammatical Error Correction
by: Kobayashi, Masamune, et al.
Published: (2024)
by: Kobayashi, Masamune, et al.
Published: (2024)
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
by: Hong, Zhaochen, et al.
Published: (2025)
by: Hong, Zhaochen, et al.
Published: (2025)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
by: Su, Jiamin, et al.
Published: (2025)
by: Su, Jiamin, et al.
Published: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
by: Holtermann, Carolin, et al.
Published: (2024)
by: Holtermann, Carolin, et al.
Published: (2024)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)
by: Leng, Jixuan, et al.
Published: (2025)
Similar Items
-
Pruning Multilingual Large Language Models for Multilingual Inference
by: Kim, Hwichan, et al.
Published: (2024) -
Oogiri-Master: Benchmarking Humor Understanding via Oogiri
by: Murakami, Soichiro, et al.
Published: (2025) -
Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
by: Zhang, Huaying, et al.
Published: (2025) -
Exploring the Effects of Alignment on Numerical Bias in Large Language Models
by: Sato, Ayako, et al.
Published: (2026) -
Assessing the Capability of LLMs in Solving POSCOMP Questions
by: Viegas, Cayo, et al.
Published: (2025)