Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914896026796032 |
|---|---|
| author | Sarti, Gabriele Caselli, Tommaso Nissim, Malvina Bisazza, Arianna |
| author_facet | Sarti, Gabriele Caselli, Tommaso Nissim, Malvina Bisazza, Arianna |
| contents | Rebuses are puzzles requiring constrained multi-step reasoning to identify a hidden phrase from a set of images and letters. In this work, we introduce a large collection of verbalized rebuses for the Italian language and use it to assess the rebus-solving capabilities of state-of-the-art large language models. While general-purpose systems such as LLaMA-3 and GPT-4o perform poorly on this task, ad-hoc fine-tuning seems to improve models' performance. However, we find that performance gains from training are largely motivated by memorization. Our results suggest that rebus solving remains a challenging test bed to evaluate large language models' linguistic proficiency and sequential instruction-following skills. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_00584 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses Sarti, Gabriele Caselli, Tommaso Nissim, Malvina Bisazza, Arianna Computation and Language Artificial Intelligence Rebuses are puzzles requiring constrained multi-step reasoning to identify a hidden phrase from a set of images and letters. In this work, we introduce a large collection of verbalized rebuses for the Italian language and use it to assess the rebus-solving capabilities of state-of-the-art large language models. While general-purpose systems such as LLaMA-3 and GPT-4o perform poorly on this task, ad-hoc fine-tuning seems to improve models' performance. However, we find that performance gains from training are largely motivated by memorization. Our results suggest that rebus solving remains a challenging test bed to evaluate large language models' linguistic proficiency and sequential instruction-following skills. |
| title | Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2408.00584 |