Limited Linguistic Diversity in Embodied AI Datasets

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wanna, Selma, Luhtaru, Agnes, Salfity, Jonathan, Barron, Ryan, Moore, Juston, Matuszek, Cynthia, Pryor, Mitch
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914514138562560
author Wanna, Selma
Luhtaru, Agnes
Salfity, Jonathan
Barron, Ryan
Moore, Juston
Matuszek, Cynthia
Pryor, Mitch
author_facet Wanna, Selma
Luhtaru, Agnes
Salfity, Jonathan
Barron, Ryan
Moore, Juston
Matuszek, Cynthia
Pryor, Mitch
contents Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate these systems remain poorly documented. In this work, we present a systematic dataset audit of several widely used VLA corpora, aiming to characterize what kinds of instructions these datasets actually contain and how much linguistic variety they provide. We quantify instruction language along complementary dimensions--including lexical variety, duplication and overlap, semantic similarity, and syntactic complexity. Our analysis shows that many datasets rely on highly repetitive, template-like commands with limited structural variation, yielding a narrow distribution of instruction forms. We position these findings as descriptive documentation of the language signal available in current VLA training and evaluation data, intended to support more detailed dataset reporting, more principled dataset selection, and targeted curation or augmentation strategies that broaden language coverage.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03136
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Limited Linguistic Diversity in Embodied AI Datasets
Wanna, Selma
Luhtaru, Agnes
Salfity, Jonathan
Barron, Ryan
Moore, Juston
Matuszek, Cynthia
Pryor, Mitch
Computation and Language
Artificial Intelligence
Robotics
Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate these systems remain poorly documented. In this work, we present a systematic dataset audit of several widely used VLA corpora, aiming to characterize what kinds of instructions these datasets actually contain and how much linguistic variety they provide. We quantify instruction language along complementary dimensions--including lexical variety, duplication and overlap, semantic similarity, and syntactic complexity. Our analysis shows that many datasets rely on highly repetitive, template-like commands with limited structural variation, yielding a narrow distribution of instruction forms. We position these findings as descriptive documentation of the language signal available in current VLA training and evaluation data, intended to support more detailed dataset reporting, more principled dataset selection, and targeted curation or augmentation strategies that broaden language coverage.
title Limited Linguistic Diversity in Embodied AI Datasets
topic Computation and Language
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2601.03136