Common Ground Tracking in Multimodal Dialogue

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khebour, Ibrahim, Lai, Kenneth, Bradford, Mariah, Zhu, Yifan, Brutti, Richard, Tam, Christopher, Tu, Jingxuan, Ibarra, Benjamin, Blanchard, Nathaniel, Krishnaswamy, Nikhil, Pustejovsky, James
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911813394759680
author Khebour, Ibrahim
Lai, Kenneth
Bradford, Mariah
Zhu, Yifan
Brutti, Richard
Tam, Christopher
Tu, Jingxuan
Ibarra, Benjamin
Blanchard, Nathaniel
Krishnaswamy, Nikhil
Pustejovsky, James
author_facet Khebour, Ibrahim
Lai, Kenneth
Bradford, Mariah
Zhu, Yifan
Brutti, Richard
Tam, Christopher
Tu, Jingxuan
Ibarra, Benjamin
Blanchard, Nathaniel
Krishnaswamy, Nikhil
Pustejovsky, James
contents Within Dialogue Modeling research in AI and NLP, considerable attention has been spent on ``dialogue state tracking'' (DST), which is the ability to update the representations of the speaker's needs at each turn in the dialogue by taking into account the past dialogue moves and history. Less studied but just as important to dialogue modeling, however, is ``common ground tracking'' (CGT), which identifies the shared belief space held by all of the participants in a task-oriented dialogue: the task-relevant propositions all participants accept as true. In this paper we present a method for automatically identifying the current set of shared beliefs and ``questions under discussion'' (QUDs) of a group with a shared goal. We annotate a dataset of multimodal interactions in a shared physical space with speech transcriptions, prosodic features, gestures, actions, and facets of collaboration, and operationalize these features for use in a deep neural model to predict moves toward construction of common ground. Model outputs cascade into a set of formal closure rules derived from situated evidence and belief axioms and update operations. We empirically assess the contribution of each feature type toward successful construction of common ground relative to ground truth, establishing a benchmark in this novel, challenging task.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17284
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Common Ground Tracking in Multimodal Dialogue
Khebour, Ibrahim
Lai, Kenneth
Bradford, Mariah
Zhu, Yifan
Brutti, Richard
Tam, Christopher
Tu, Jingxuan
Ibarra, Benjamin
Blanchard, Nathaniel
Krishnaswamy, Nikhil
Pustejovsky, James
Computation and Language
Within Dialogue Modeling research in AI and NLP, considerable attention has been spent on ``dialogue state tracking'' (DST), which is the ability to update the representations of the speaker's needs at each turn in the dialogue by taking into account the past dialogue moves and history. Less studied but just as important to dialogue modeling, however, is ``common ground tracking'' (CGT), which identifies the shared belief space held by all of the participants in a task-oriented dialogue: the task-relevant propositions all participants accept as true. In this paper we present a method for automatically identifying the current set of shared beliefs and ``questions under discussion'' (QUDs) of a group with a shared goal. We annotate a dataset of multimodal interactions in a shared physical space with speech transcriptions, prosodic features, gestures, actions, and facets of collaboration, and operationalize these features for use in a deep neural model to predict moves toward construction of common ground. Model outputs cascade into a set of formal closure rules derived from situated evidence and belief axioms and update operations. We empirically assess the contribution of each feature type toward successful construction of common ground relative to ground truth, establishing a benchmark in this novel, challenging task.
title Common Ground Tracking in Multimodal Dialogue
topic Computation and Language
url https://arxiv.org/abs/2403.17284