Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908587482152960 |
|---|---|
| author | Wang, Yufei Kovashka, Adriana Fernández, Loretta Coutanche, Marc N. Wiener, Seth |
| author_facet | Wang, Yufei Kovashka, Adriana Fernández, Loretta Coutanche, Marc N. Wiener, Seth |
| contents | We investigate a new setting for foreign language learning, where learners infer the meaning of unfamiliar words in a multimodal context of a sentence describing a paired image. We conduct studies with human participants using different image-text pairs. We analyze the features of the data (i.e., images and texts) that make it easier for participants to infer the meaning of a masked or unfamiliar word, and what language backgrounds of the participants correlate with success. We find only some intuitive features have strong correlations with participant performance, prompting the need for further investigating of predictive features for success in these tasks. We also analyze the ability of AI systems to reason about participant performance, and discover promising future directions for improving this reasoning ability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_09815 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning Wang, Yufei Kovashka, Adriana Fernández, Loretta Coutanche, Marc N. Wiener, Seth Computer Vision and Pattern Recognition Artificial Intelligence We investigate a new setting for foreign language learning, where learners infer the meaning of unfamiliar words in a multimodal context of a sentence describing a paired image. We conduct studies with human participants using different image-text pairs. We analyze the features of the data (i.e., images and texts) that make it easier for participants to infer the meaning of a masked or unfamiliar word, and what language backgrounds of the participants correlate with success. We find only some intuitive features have strong correlations with participant performance, prompting the need for further investigating of predictive features for success in these tasks. We also analyze the ability of AI systems to reason about participant performance, and discover promising future directions for improving this reasoning ability. |
| title | Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2510.09815 |