Words at Play: Benchmarking Audio Pun Understanding in Large Audio-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Yuchen, Zhong, Shaoxin, Zhu, Yonghua, Wang, Ruofan, Huang, Zijian, Wang, Qiqi, Zhao, Na, Benavides-Prado, Diana, Witbrock, Michael
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915874522267648
author Su, Yuchen
Zhong, Shaoxin
Zhu, Yonghua
Wang, Ruofan
Huang, Zijian
Wang, Qiqi
Zhao, Na
Benavides-Prado, Diana
Witbrock, Michael
author_facet Su, Yuchen
Zhong, Shaoxin
Zhu, Yonghua
Wang, Ruofan
Huang, Zijian
Wang, Qiqi
Zhao, Na
Benavides-Prado, Diana
Witbrock, Michael
contents Puns represent a typical linguistic phenomenon that exploits polysemy and phonetic ambiguity to generate humour, posing unique challenges for natural language understanding. Within pun research, audio plays a central role in human communication except text and images, while datasets and systematic resources for spoken puns remain scarce, leaving this crucial modality largely underexplored. In this paper, we present APUN-Bench, the first benchmark dedicated to evaluating large audio language models (LALMs) on audio pun understanding. Our benchmark contains 4,434 audio samples annotated across three stages: pun recognition, pun word location and pun meaning inference. We conduct a deep analysis of APUN-Bench by systematically evaluating 10 state-of-the-art LALMs, uncovering substantial performance gaps in recognizing, localizing, and interpreting audio puns. This analysis reveals key challenges, such as positional biases in audio pun location and error cases in meaning inference, offering actionable insights for advancing humour-aware audio intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18678
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Words at Play: Benchmarking Audio Pun Understanding in Large Audio-Language Models
Su, Yuchen
Zhong, Shaoxin
Zhu, Yonghua
Wang, Ruofan
Huang, Zijian
Wang, Qiqi
Zhao, Na
Benavides-Prado, Diana
Witbrock, Michael
Sound
Computation and Language
Puns represent a typical linguistic phenomenon that exploits polysemy and phonetic ambiguity to generate humour, posing unique challenges for natural language understanding. Within pun research, audio plays a central role in human communication except text and images, while datasets and systematic resources for spoken puns remain scarce, leaving this crucial modality largely underexplored. In this paper, we present APUN-Bench, the first benchmark dedicated to evaluating large audio language models (LALMs) on audio pun understanding. Our benchmark contains 4,434 audio samples annotated across three stages: pun recognition, pun word location and pun meaning inference. We conduct a deep analysis of APUN-Bench by systematically evaluating 10 state-of-the-art LALMs, uncovering substantial performance gaps in recognizing, localizing, and interpreting audio puns. This analysis reveals key challenges, such as positional biases in audio pun location and error cases in meaning inference, offering actionable insights for advancing humour-aware audio intelligence.
title Words at Play: Benchmarking Audio Pun Understanding in Large Audio-Language Models
topic Sound
Computation and Language
url https://arxiv.org/abs/2603.18678