Robust Pitch Estimation and Tracking for Speakers Based on Subband Encoding and the Generalized Labeled Multi-Bernoulli Filter
Fuente:
arXiv
Guardado en:
| Autor principal: | Lin, Shoufeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reverberation-Robust Localization of Speakers Using Distinct Speech Onsets and Multi-channel Cross-Correlations
por: Lin, Shoufeng
Publicado: (2026)
por: Lin, Shoufeng
Publicado: (2026)
A Generalized Weighted Overlap-Add (WOLA) Filter Bank for Improved Subband System Identification
por: Sharma, Mohit, et al.
Publicado: (2025)
por: Sharma, Mohit, et al.
Publicado: (2025)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
por: Rautenberg, Frederik, et al.
Publicado: (2026)
por: Rautenberg, Frederik, et al.
Publicado: (2026)
Nearest Kronecker Product Decomposition Based Subband Adaptive Filter: Algorithms and Applications
por: Ye, Jianhong, et al.
Publicado: (2026)
por: Ye, Jianhong, et al.
Publicado: (2026)
Robust Training for Speaker Verification against Noisy Labels
por: Fang, Zhihua, et al.
Publicado: (2022)
por: Fang, Zhihua, et al.
Publicado: (2022)
Harmonic Summation-Based Robust Pitch Estimation in Noisy and Reverberant Environments
por: Singh, Anup, et al.
Publicado: (2025)
por: Singh, Anup, et al.
Publicado: (2025)
Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception
por: Chen, Jiaxin, et al.
Publicado: (2025)
por: Chen, Jiaxin, et al.
Publicado: (2025)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
por: Terashima, Ryo, et al.
Publicado: (2025)
por: Terashima, Ryo, et al.
Publicado: (2025)
What Does the Speaker Embedding Encode?
por: Wang, Shuai, et al.
Publicado: (2025)
por: Wang, Shuai, et al.
Publicado: (2025)
Fractional-Order Subband p-Norm Adaptive Filter via Transformation Nearest Kronecker Product Decomposition for Active Noise Control
por: Ye, Jianhong, et al.
Publicado: (2026)
por: Ye, Jianhong, et al.
Publicado: (2026)
Subband Architecture Aided Selective Fixed-Filter Active Noise Control
por: Liang, Hong-Cheng, et al.
Publicado: (2025)
por: Liang, Hong-Cheng, et al.
Publicado: (2025)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
por: Wei, Haojie, et al.
Publicado: (2023)
por: Wei, Haojie, et al.
Publicado: (2023)
Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone Array
por: Qiao, Yue, et al.
Publicado: (2024)
por: Qiao, Yue, et al.
Publicado: (2024)
Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity
por: Subramani, Krishna, et al.
Publicado: (2023)
por: Subramani, Krishna, et al.
Publicado: (2023)
A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation
por: Taenzer, Michael
Publicado: (2026)
por: Taenzer, Michael
Publicado: (2026)
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
por: Chung, Woo-Jin, et al.
Publicado: (2023)
por: Chung, Woo-Jin, et al.
Publicado: (2023)
Neural Forward Filtering for Speaker-Image Separation
por: Sun, Jingqi, et al.
Publicado: (2025)
por: Sun, Jingqi, et al.
Publicado: (2025)
Multi-Label Training for Text-Independent Speaker Identification
por: Xue, Yuqi
Publicado: (2022)
por: Xue, Yuqi
Publicado: (2022)
Robust Target Speaker Direction of Arrival Estimation
por: Li, Zixuan, et al.
Publicado: (2024)
por: Li, Zixuan, et al.
Publicado: (2024)
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
por: Cheng, Ming, et al.
Publicado: (2024)
por: Cheng, Ming, et al.
Publicado: (2024)
Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
por: Bhattacharjee, Sankha Subhra, et al.
Publicado: (2024)
por: Bhattacharjee, Sankha Subhra, et al.
Publicado: (2024)
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
por: Dai, Yuhang, et al.
Publicado: (2026)
por: Dai, Yuhang, et al.
Publicado: (2026)
Cross-domain Neural Pitch and Periodicity Estimation
por: Morrison, Max, et al.
Publicado: (2023)
por: Morrison, Max, et al.
Publicado: (2023)
Improving Neural Pitch Estimation with SWIPE Kernels
por: Marttila, David, et al.
Publicado: (2025)
por: Marttila, David, et al.
Publicado: (2025)
Meta-Learning-Based Delayless Subband Adaptive Filter using Complex Self-Attention for Active Noise Control
por: Feng, Pengxing, et al.
Publicado: (2024)
por: Feng, Pengxing, et al.
Publicado: (2024)
IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
por: Liu, Zeyan, et al.
Publicado: (2025)
por: Liu, Zeyan, et al.
Publicado: (2025)
Joint Optimization of Speaker and Spoof Detectors for Spoofing-Robust Automatic Speaker Verification
por: Kurnaz, Oğuzhan, et al.
Publicado: (2025)
por: Kurnaz, Oğuzhan, et al.
Publicado: (2025)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
por: Shao, Yiwen, et al.
Publicado: (2024)
por: Shao, Yiwen, et al.
Publicado: (2024)
Flexible Multi-Channel Target Speaker Extraction Using Geometry-Conditioned Spatially Selective Non-linear Filters
por: Li, Jiatong, et al.
Publicado: (2026)
por: Li, Jiatong, et al.
Publicado: (2026)
A Robust Method for Pitch Tracking in the Frequency Following Response using Harmonic Amplitude Summation Filterbank
por: Sadeghkhani, Sajad, et al.
Publicado: (2025)
por: Sadeghkhani, Sajad, et al.
Publicado: (2025)
Multi-View Based Audio Visual Target Speaker Extraction
por: Yang, Peijun, et al.
Publicado: (2026)
por: Yang, Peijun, et al.
Publicado: (2026)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
por: Zhang, Lin, et al.
Publicado: (2024)
por: Zhang, Lin, et al.
Publicado: (2024)
An Age-Agnostic System for Robust Speaker Verification
por: Zheng, Jiusi, et al.
Publicado: (2025)
por: Zheng, Jiusi, et al.
Publicado: (2025)
Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification
por: Gan, Chong-Xin, et al.
Publicado: (2023)
por: Gan, Chong-Xin, et al.
Publicado: (2023)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
por: Riou, Alain, et al.
Publicado: (2023)
por: Riou, Alain, et al.
Publicado: (2023)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Toward Fully Self-Supervised Multi-Pitch Estimation
por: Cwitkowitz, Frank, et al.
Publicado: (2024)
por: Cwitkowitz, Frank, et al.
Publicado: (2024)
Multi-Speaker DOA Estimation in Binaural Hearing Aids using Deep Learning and Speaker Count Fusion
por: Jazaeri, Farnaz, et al.
Publicado: (2025)
por: Jazaeri, Farnaz, et al.
Publicado: (2025)
UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
por: Gan, Chong-Xin, et al.
Publicado: (2026)
por: Gan, Chong-Xin, et al.
Publicado: (2026)
Ejemplares similares
-
Reverberation-Robust Localization of Speakers Using Distinct Speech Onsets and Multi-channel Cross-Correlations
por: Lin, Shoufeng
Publicado: (2026) -
A Generalized Weighted Overlap-Add (WOLA) Filter Bank for Improved Subband System Identification
por: Sharma, Mohit, et al.
Publicado: (2025) -
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
por: Rautenberg, Frederik, et al.
Publicado: (2026) -
Nearest Kronecker Product Decomposition Based Subband Adaptive Filter: Algorithms and Applications
por: Ye, Jianhong, et al.
Publicado: (2026) -
Robust Training for Speaker Verification against Noisy Labels
por: Fang, Zhihua, et al.
Publicado: (2022)