Public speech recognition transcripts as a configuring parameter

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rudaz, Damien, Licoppe, Christian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915231415926784
author Rudaz, Damien
Licoppe, Christian
author_facet Rudaz, Damien
Licoppe, Christian
contents Displaying a written transcript of what a human said (i.e. producing an "automatic speech recognition transcript") is a common feature for smartphone vocal assistants: the utterance produced by a human speaker (e.g. a question) is displayed on the screen while it is being verbally responded to by the vocal assistant. Although very rarely, this feature also exists on some "social" robots which transcribe human interactants' speech on a screen or a tablet. We argue that this informational configuration is pragmatically consequential on the interaction, both for human participants and for the embodied conversational agent. Based on a corpus of co-present interactions with a humanoid robot, we attempt to show that this transcript is a contextual feature which can heavily impact the actions ascribed by humans to the robot: that is, the way in which humans respond to the robot's behavior as constituting a specific type of action (rather than another) and as constituting an adequate response to their own previous turn.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04488
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Public speech recognition transcripts as a configuring parameter
Rudaz, Damien
Licoppe, Christian
Human-Computer Interaction
Displaying a written transcript of what a human said (i.e. producing an "automatic speech recognition transcript") is a common feature for smartphone vocal assistants: the utterance produced by a human speaker (e.g. a question) is displayed on the screen while it is being verbally responded to by the vocal assistant. Although very rarely, this feature also exists on some "social" robots which transcribe human interactants' speech on a screen or a tablet. We argue that this informational configuration is pragmatically consequential on the interaction, both for human participants and for the embodied conversational agent. Based on a corpus of co-present interactions with a humanoid robot, we attempt to show that this transcript is a contextual feature which can heavily impact the actions ascribed by humans to the robot: that is, the way in which humans respond to the robot's behavior as constituting a specific type of action (rather than another) and as constituting an adequate response to their own previous turn.
title Public speech recognition transcripts as a configuring parameter
topic Human-Computer Interaction
url https://arxiv.org/abs/2504.04488