Effective Context in Neural Speech Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Yen, Goldwater, Sharon, Tang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913864191311872
author Meng, Yen
Goldwater, Sharon
Tang, Hao
author_facet Meng, Yen
Goldwater, Sharon
Tang, Hao
contents Modern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e., the effective context. Here, we propose two approaches to measuring the effective context, and use them to analyze different speech Transformers. For supervised models, we find that the effective context correlates well with the nature of the task, with fundamental frequency tracking, phone classification, and word classification requiring increasing amounts of effective context. For self-supervised models, we find that effective context increases mainly in the early layers, and remains relatively short -- similar to the supervised phone model. Given that these models do not use a long context during prediction, we show that HuBERT can be run in streaming mode without modification to the architecture and without further fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22487
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Effective Context in Neural Speech Models
Meng, Yen
Goldwater, Sharon
Tang, Hao
Sound
Computation and Language
Audio and Speech Processing
Modern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e., the effective context. Here, we propose two approaches to measuring the effective context, and use them to analyze different speech Transformers. For supervised models, we find that the effective context correlates well with the nature of the task, with fundamental frequency tracking, phone classification, and word classification requiring increasing amounts of effective context. For self-supervised models, we find that effective context increases mainly in the early layers, and remains relatively short -- similar to the supervised phone model. Given that these models do not use a long context during prediction, we show that HuBERT can be run in streaming mode without modification to the architecture and without further fine-tuning.
title Effective Context in Neural Speech Models
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2505.22487