Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Samo, Giuseppe, Merlo, Paola
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912968452603904
author Samo, Giuseppe
Merlo, Paola
author_facet Samo, Giuseppe
Merlo, Paola
contents Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this work, we present curated paradigm-based datasets for four languages, designed to probe systematic cross-sentence knowledge of verb alternations (change-of-state and object-drop constructions in English, German and Italian, and Hebrew binyanim). The datasets comprise thousands of the Blackbird Language Matrices (BLMs) problems. The BLM task -- an RPM/ARC-like task devised specifically for language -- is a controlled linguistic puzzle where models must select the sentence that completes a pattern according to syntactic and semantic rules. We introduce three types of templates varying in complexity and apply linguistically-informed data augmentation strategies across synthetic and natural data. We provide simple baseline performance results across English, Italian, German, and Hebrew, that demonstrate the diagnostic usefulness of the datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15295
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies
Samo, Giuseppe
Merlo, Paola
Computation and Language
Databases
Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this work, we present curated paradigm-based datasets for four languages, designed to probe systematic cross-sentence knowledge of verb alternations (change-of-state and object-drop constructions in English, German and Italian, and Hebrew binyanim). The datasets comprise thousands of the Blackbird Language Matrices (BLMs) problems. The BLM task -- an RPM/ARC-like task devised specifically for language -- is a controlled linguistic puzzle where models must select the sentence that completes a pattern according to syntactic and semantic rules. We introduce three types of templates varying in complexity and apply linguistically-informed data augmentation strategies across synthetic and natural data. We provide simple baseline performance results across English, Italian, German, and Hebrew, that demonstrate the diagnostic usefulness of the datasets.
title Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies
topic Computation and Language
Databases
url https://arxiv.org/abs/2603.15295