Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dong, Dota Tianai, Luo, Yifan, Wang, Po-Ya Angela, Ozyurek, Asli, Rubio-Fernandez, Paula
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913046665887744
author Dong, Dota Tianai
Luo, Yifan
Wang, Po-Ya Angela
Ozyurek, Asli
Rubio-Fernandez, Paula
author_facet Dong, Dota Tianai
Luo, Yifan
Wang, Po-Ya Angela
Ozyurek, Asli
Rubio-Fernandez, Paula
contents Multimodal language models (MLMs) increasingly demonstrate human-like communication, yet their use of everyday perspectival words remains poorly understood. To address this gap, we compare humans and MLMs in their use of three word types that impose increasing cognitive demands: vocabulary (for example, "boat" or "cup"), possessives (for example, "mine" versus "yours"), and demonstratives (for example, "this one" versus "that one"). Testing seven MLMs against human participants, we find that perspectival words are harder than vocabulary words for both groups. The gap is larger for MLMs: while models approach human-level performance on vocabulary, they show clear deficits with possessives and even greater difficulty with demonstratives. Ablation analyses indicate that limitations in perspective-taking and spatial reasoning are key sources of these gaps. Instruction-based prompting reduces the gap for possessives but leaves demonstratives far below human performance. These results show that, unlike vocabulary, perspectival words pose a greater challenge in human communication, and this difficulty is amplified in MLMs, revealing a shortfall in their pragmatic and social-cognitive abilities.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00065
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models
Dong, Dota Tianai
Luo, Yifan
Wang, Po-Ya Angela
Ozyurek, Asli
Rubio-Fernandez, Paula
Computation and Language
Artificial Intelligence
Multimodal language models (MLMs) increasingly demonstrate human-like communication, yet their use of everyday perspectival words remains poorly understood. To address this gap, we compare humans and MLMs in their use of three word types that impose increasing cognitive demands: vocabulary (for example, "boat" or "cup"), possessives (for example, "mine" versus "yours"), and demonstratives (for example, "this one" versus "that one"). Testing seven MLMs against human participants, we find that perspectival words are harder than vocabulary words for both groups. The gap is larger for MLMs: while models approach human-level performance on vocabulary, they show clear deficits with possessives and even greater difficulty with demonstratives. Ablation analyses indicate that limitations in perspective-taking and spatial reasoning are key sources of these gaps. Instruction-based prompting reduces the gap for possessives but leaves demonstratives far below human performance. These results show that, unlike vocabulary, perspectival words pose a greater challenge in human communication, and this difficulty is amplified in MLMs, revealing a shortfall in their pragmatic and social-cognitive abilities.
title Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.00065