Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koguchi, Junya, Koriyama, Tomoki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917241400852480
author Koguchi, Junya
Koriyama, Tomoki
author_facet Koguchi, Junya
Koriyama, Tomoki
contents The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a principled analysis and improvement of this technique. First, we offer a theoretical basis for its effectiveness, explaining the error variance reduction for fundamental frequency estimation and invoking Condorcet's jury theorem for voiced/unvoiced detection accuracy. To address its practical limitations, we propose two key improvements: 1) a pre-voting alignment procedure to correct temporal and frequential biases among estimators, and 2) a greedy algorithm to select a compact yet effective subset of estimators based on error correlation. Experiments on a diverse dataset of speech, singing, and music show that our proposed method with alignment outperforms individual state-of-the-art estimators in clean conditions and maintains robust voiced/unvoiced detection in noisy environments.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01727
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection
Koguchi, Junya
Koriyama, Tomoki
Sound
The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a principled analysis and improvement of this technique. First, we offer a theoretical basis for its effectiveness, explaining the error variance reduction for fundamental frequency estimation and invoking Condorcet's jury theorem for voiced/unvoiced detection accuracy. To address its practical limitations, we propose two key improvements: 1) a pre-voting alignment procedure to correct temporal and frequential biases among estimators, and 2) a greedy algorithm to select a compact yet effective subset of estimators based on error correlation. Experiments on a diverse dataset of speech, singing, and music show that our proposed method with alignment outperforms individual state-of-the-art estimators in clean conditions and maintains robust voiced/unvoiced detection in noisy environments.
title Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection
topic Sound
url https://arxiv.org/abs/2602.01727