Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Samo, Giuseppe, Merlo, Paola
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915892471791616
author Samo, Giuseppe
Merlo, Paola
author_facet Samo, Giuseppe
Merlo, Paola
contents This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (BLMs), structured datasets designed to probe linguistic knowledge of underlying patterns across sentence sets. We compare structured templates instantiated with natural sentences extracted from Universal Dependencies to structured templates of synthetic sentences. Experiments show that while models achieve ceiling performance when trained and tested on synthetic datasets, they do not reliably generalize to natural sentences. In contrast, models trained on natural data exhibit robust performance across both natural and synthetic test suites, demonstrating their superior ability to capture abstract linguistic patterns. These results corroborate the value of natural data and of structured set ups in linguistic evaluation for probing LLMs' syntactic and semantic knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25227
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian
Samo, Giuseppe
Merlo, Paola
Computation and Language
This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (BLMs), structured datasets designed to probe linguistic knowledge of underlying patterns across sentence sets. We compare structured templates instantiated with natural sentences extracted from Universal Dependencies to structured templates of synthetic sentences. Experiments show that while models achieve ceiling performance when trained and tested on synthetic datasets, they do not reliably generalize to natural sentences. In contrast, models trained on natural data exhibit robust performance across both natural and synthetic test suites, demonstrating their superior ability to capture abstract linguistic patterns. These results corroborate the value of natural data and of structured set ups in linguistic evaluation for probing LLMs' syntactic and semantic knowledge.
title Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian
topic Computation and Language
url https://arxiv.org/abs/2603.25227