Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909780350599168 |
|---|---|
| author | Bovo, Riccardo Brudy, Frederik Fitzmaurice, George Anderson, Fraser |
| author_facet | Bovo, Riccardo Brudy, Frederik Fitzmaurice, George Anderson, Fraser |
| contents | Understanding transcripts of immersive multimodal conversations is challenging because speakers frequently rely on visual context and non-verbal cues, such as gestures and visual attention, which are not captured in speech alone. This lack of information makes coreferences resolution-the task of linking ambiguous expressions like ``it'' or ``there'' to their intended referents-particularly challenging. In this paper we present a system that augments VR speech transcript with eye-tracking laser pointing data, and scene metadata to generate textual descriptions of non-verbal communication and the corresponding objects of interest. To evaluate the system, we collected gaze, gesture, and voice data from 12 participants (6 pairs) engaged in an open-ended design critique of a 3D model of an apartment. Our results show a 26.5\% improvement in coreference resolution accuracy by a GPT model when using our multimodal transcript compared to a speech-only baseline. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_08689 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution Bovo, Riccardo Brudy, Frederik Fitzmaurice, George Anderson, Fraser Human-Computer Interaction Understanding transcripts of immersive multimodal conversations is challenging because speakers frequently rely on visual context and non-verbal cues, such as gestures and visual attention, which are not captured in speech alone. This lack of information makes coreferences resolution-the task of linking ambiguous expressions like ``it'' or ``there'' to their intended referents-particularly challenging. In this paper we present a system that augments VR speech transcript with eye-tracking laser pointing data, and scene metadata to generate textual descriptions of non-verbal communication and the corresponding objects of interest. To evaluate the system, we collected gaze, gesture, and voice data from 12 participants (6 pairs) engaged in an open-ended design critique of a 3D model of an apartment. Our results show a 26.5\% improvement in coreference resolution accuracy by a GPT model when using our multimodal transcript compared to a speech-only baseline. |
| title | Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution |
| topic | Human-Computer Interaction |
| url | https://arxiv.org/abs/2509.08689 |