Saved in:
Bibliographic Details
Main Authors: Yadavalli, Aditya, Pimentel, Tiago, Regev, Tamar I, Wilcox, Ethan, Warstadt, Alex
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.16832
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918255401107456
author Yadavalli, Aditya
Pimentel, Tiago
Regev, Tamar I
Wilcox, Ethan
Warstadt, Alex
author_facet Yadavalli, Aditya
Pimentel, Tiago
Regev, Tamar I
Wilcox, Ethan
Warstadt, Alex
contents Prosody -- the melody of speech -- conveys critical information often not captured by the words or text of a message. In this paper, we propose an information-theoretic approach to quantify how much information is expressed by prosody alone and not by text, and crucially, what that information is about. Our approach applies large speech and language models to estimate the mutual information between a particular dimension of an utterance's meaning (e.g., its emotion) and any of its communication channels (e.g., audio or text). We then use this approach to quantify how much information is conveyed by audio and text about sarcasm, emotion, and questionhood, using speech from television and podcasts. We find that for sarcasm and emotion the audio channel -- and by implication the prosodic channel -- transmits over an order of magnitude more information about these features than the text channel alone, at least when long-term context beyond the current sentence is unavailable. For questionhood, prosody provides comparatively less additional information. We conclude by outlining a program applying our approach to more dimensions of meaning, communication channels, and languages.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
Yadavalli, Aditya
Pimentel, Tiago
Regev, Tamar I
Wilcox, Ethan
Warstadt, Alex
Computation and Language
Prosody -- the melody of speech -- conveys critical information often not captured by the words or text of a message. In this paper, we propose an information-theoretic approach to quantify how much information is expressed by prosody alone and not by text, and crucially, what that information is about. Our approach applies large speech and language models to estimate the mutual information between a particular dimension of an utterance's meaning (e.g., its emotion) and any of its communication channels (e.g., audio or text). We then use this approach to quantify how much information is conveyed by audio and text about sarcasm, emotion, and questionhood, using speech from television and podcasts. We find that for sarcasm and emotion the audio channel -- and by implication the prosodic channel -- transmits over an order of magnitude more information about these features than the text channel alone, at least when long-term context beyond the current sentence is unavailable. For questionhood, prosody provides comparatively less additional information. We conclude by outlining a program applying our approach to more dimensions of meaning, communication channels, and languages.
title What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
topic Computation and Language
url https://arxiv.org/abs/2512.16832