Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sarti, Gabriele, Caselli, Tommaso, Nissim, Malvina, Bisazza, Arianna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914896026796032
author Sarti, Gabriele
Caselli, Tommaso
Nissim, Malvina
Bisazza, Arianna
author_facet Sarti, Gabriele
Caselli, Tommaso
Nissim, Malvina
Bisazza, Arianna
contents Rebuses are puzzles requiring constrained multi-step reasoning to identify a hidden phrase from a set of images and letters. In this work, we introduce a large collection of verbalized rebuses for the Italian language and use it to assess the rebus-solving capabilities of state-of-the-art large language models. While general-purpose systems such as LLaMA-3 and GPT-4o perform poorly on this task, ad-hoc fine-tuning seems to improve models' performance. However, we find that performance gains from training are largely motivated by memorization. Our results suggest that rebus solving remains a challenging test bed to evaluate large language models' linguistic proficiency and sequential instruction-following skills.
format Preprint
id arxiv_https___arxiv_org_abs_2408_00584
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses
Sarti, Gabriele
Caselli, Tommaso
Nissim, Malvina
Bisazza, Arianna
Computation and Language
Artificial Intelligence
Rebuses are puzzles requiring constrained multi-step reasoning to identify a hidden phrase from a set of images and letters. In this work, we introduce a large collection of verbalized rebuses for the Italian language and use it to assess the rebus-solving capabilities of state-of-the-art large language models. While general-purpose systems such as LLaMA-3 and GPT-4o perform poorly on this task, ad-hoc fine-tuning seems to improve models' performance. However, we find that performance gains from training are largely motivated by memorization. Our results suggest that rebus solving remains a challenging test bed to evaluate large language models' linguistic proficiency and sequential instruction-following skills.
title Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2408.00584