Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yufei, Kovashka, Adriana, Fernández, Loretta, Coutanche, Marc N., Wiener, Seth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908587482152960
author Wang, Yufei
Kovashka, Adriana
Fernández, Loretta
Coutanche, Marc N.
Wiener, Seth
author_facet Wang, Yufei
Kovashka, Adriana
Fernández, Loretta
Coutanche, Marc N.
Wiener, Seth
contents We investigate a new setting for foreign language learning, where learners infer the meaning of unfamiliar words in a multimodal context of a sentence describing a paired image. We conduct studies with human participants using different image-text pairs. We analyze the features of the data (i.e., images and texts) that make it easier for participants to infer the meaning of a masked or unfamiliar word, and what language backgrounds of the participants correlate with success. We find only some intuitive features have strong correlations with participant performance, prompting the need for further investigating of predictive features for success in these tasks. We also analyze the ability of AI systems to reason about participant performance, and discover promising future directions for improving this reasoning ability.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
Wang, Yufei
Kovashka, Adriana
Fernández, Loretta
Coutanche, Marc N.
Wiener, Seth
Computer Vision and Pattern Recognition
Artificial Intelligence
We investigate a new setting for foreign language learning, where learners infer the meaning of unfamiliar words in a multimodal context of a sentence describing a paired image. We conduct studies with human participants using different image-text pairs. We analyze the features of the data (i.e., images and texts) that make it easier for participants to infer the meaning of a masked or unfamiliar word, and what language backgrounds of the participants correlate with success. We find only some intuitive features have strong correlations with participant performance, prompting the need for further investigating of predictive features for success in these tasks. We also analyze the ability of AI systems to reason about participant performance, and discover promising future directions for improving this reasoning ability.
title Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.09815