Chumor 2.0: Towards Benchmarking Chinese Humor Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Ruiqi, He, Yushu, Bai, Longju, Liu, Jiarui, Sun, Zhenjie, Tang, Zenghao, Wang, He, Xia, Hanchen, Mihalcea, Rada, Deng, Naihao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916540170895360
author He, Ruiqi
He, Yushu
Bai, Longju
Liu, Jiarui
Sun, Zhenjie
Tang, Zenghao
Wang, He
Xia, Hanchen
Mihalcea, Rada
Deng, Naihao
author_facet He, Ruiqi
He, Yushu
Bai, Longju
Liu, Jiarui
Sun, Zhenjie
Tang, Zenghao
Wang, He
Xia, Hanchen
Mihalcea, Rada
Deng, Naihao
contents Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, we construct Chumor, the first Chinese humor explanation dataset that exceeds the size of existing humor datasets. Chumor is sourced from Ruo Zhi Ba, a Chinese Reddit-like platform known for sharing intellectually challenging and culturally specific jokes. We test ten LLMs through direct and chain-of-thought prompting, revealing that Chumor poses significant challenges to existing LLMs, with their accuracy slightly above random and far below human. In addition, our analysis highlights that human-annotated humor explanations are significantly better than those generated by GPT-4o and ERNIE-4-turbo. We release Chumor at https://huggingface.co/datasets/dnaihao/Chumor, our project page is at https://dnaihao.github.io/Chumor-dataset/, our leaderboard is at https://huggingface.co/spaces/dnaihao/Chumor, and our codebase is at https://github.com/dnaihao/Chumor-dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17729
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
He, Ruiqi
He, Yushu
Bai, Longju
Liu, Jiarui
Sun, Zhenjie
Tang, Zenghao
Wang, He
Xia, Hanchen
Mihalcea, Rada
Deng, Naihao
Computation and Language
Artificial Intelligence
Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, we construct Chumor, the first Chinese humor explanation dataset that exceeds the size of existing humor datasets. Chumor is sourced from Ruo Zhi Ba, a Chinese Reddit-like platform known for sharing intellectually challenging and culturally specific jokes. We test ten LLMs through direct and chain-of-thought prompting, revealing that Chumor poses significant challenges to existing LLMs, with their accuracy slightly above random and far below human. In addition, our analysis highlights that human-annotated humor explanations are significantly better than those generated by GPT-4o and ERNIE-4-turbo. We release Chumor at https://huggingface.co/datasets/dnaihao/Chumor, our project page is at https://dnaihao.github.io/Chumor-dataset/, our leaderboard is at https://huggingface.co/spaces/dnaihao/Chumor, and our codebase is at https://github.com/dnaihao/Chumor-dataset.
title Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.17729