VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marmor, Yanir, Zulti, Arad, Krongauz, David, Gabet, Adam, Snapir, Yoad, Lifshitz, Yair, Segal, Eran
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912944194846720
author Marmor, Yanir
Zulti, Arad
Krongauz, David
Gabet, Adam
Snapir, Yoad
Lifshitz, Yair
Segal, Eran
author_facet Marmor, Yanir
Zulti, Arad
Krongauz, David
Gabet, Adam
Snapir, Yoad
Lifshitz, Yair
Segal, Eran
contents Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary speech spanning 2009-2025, comprising 393 speakers with recording spans of up to 15 years. Each segment includes aligned transcripts and verified demographic metadata from official parliamentary records. We benchmark modern speech embeddings (WavLM-Large, ECAPA-TDNN, Wav2Vec2-XLSR-1B) on age prediction and speaker verification under longitudinal conditions. Speaker verification EER rises from 2.15\% to 4.58\% over 15 years for the strongest model, and cross-sectionally trained age regressors fail to capture within-speaker aging, while longitudinally trained models recover a meaningful temporal signal. We publicly release the dataset and pipeline to support aging-robust speech systems and Hebrew speech processing.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01270
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
Marmor, Yanir
Zulti, Arad
Krongauz, David
Gabet, Adam
Snapir, Yoad
Lifshitz, Yair
Segal, Eran
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Signal Processing
Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary speech spanning 2009-2025, comprising 393 speakers with recording spans of up to 15 years. Each segment includes aligned transcripts and verified demographic metadata from official parliamentary records. We benchmark modern speech embeddings (WavLM-Large, ECAPA-TDNN, Wav2Vec2-XLSR-1B) on age prediction and speaker verification under longitudinal conditions. Speaker verification EER rises from 2.15\% to 4.58\% over 15 years for the strongest model, and cross-sectionally trained age regressors fail to capture within-speaker aging, while longitudinally trained models recover a meaningful temporal signal. We publicly release the dataset and pipeline to support aging-robust speech systems and Hebrew speech processing.
title VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2603.01270