SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
Fuente:
arXiv
Salvato in:
| Autori principali: | Terashima, Ryo, Shirahata, Yuma, Kawamura, Masaya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
di: Ohnaka, Hien, et al.
Pubblicazione: (2026)
di: Ohnaka, Hien, et al.
Pubblicazione: (2026)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
di: Kawamura, Masaya, et al.
Pubblicazione: (2025)
di: Kawamura, Masaya, et al.
Pubblicazione: (2025)
Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity
di: Subramani, Krishna, et al.
Pubblicazione: (2023)
di: Subramani, Krishna, et al.
Pubblicazione: (2023)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024)
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024)
Toward Fully Self-Supervised Multi-Pitch Estimation
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2024)
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2024)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
di: Kawamura, Masaya, et al.
Pubblicazione: (2024)
di: Kawamura, Masaya, et al.
Pubblicazione: (2024)
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2025)
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2025)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
di: Riou, Alain, et al.
Pubblicazione: (2023)
di: Riou, Alain, et al.
Pubblicazione: (2023)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
di: Rautenberg, Frederik, et al.
Pubblicazione: (2026)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2026)
CC-G2PnP: Streaming Grapheme-to-Phoneme and prosody with Conformer-CTC for unsegmented languages
di: Shirahata, Yuma, et al.
Pubblicazione: (2026)
di: Shirahata, Yuma, et al.
Pubblicazione: (2026)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
di: Yoneyama, Reo, et al.
Pubblicazione: (2025)
di: Yoneyama, Reo, et al.
Pubblicazione: (2025)
Improving Neural Pitch Estimation with SWIPE Kernels
di: Marttila, David, et al.
Pubblicazione: (2025)
di: Marttila, David, et al.
Pubblicazione: (2025)
Cross-domain Neural Pitch and Periodicity Estimation
di: Morrison, Max, et al.
Pubblicazione: (2023)
di: Morrison, Max, et al.
Pubblicazione: (2023)
Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
Universal Score-based Speech Enhancement with High Content Preservation
di: Scheibler, Robin, et al.
Pubblicazione: (2024)
di: Scheibler, Robin, et al.
Pubblicazione: (2024)
Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
di: Valin, Jean-Marc, et al.
Pubblicazione: (2024)
di: Valin, Jean-Marc, et al.
Pubblicazione: (2024)
Harmonic Summation-Based Robust Pitch Estimation in Noisy and Reverberant Environments
di: Singh, Anup, et al.
Pubblicazione: (2025)
di: Singh, Anup, et al.
Pubblicazione: (2025)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
di: Wei, Haojie, et al.
Pubblicazione: (2023)
di: Wei, Haojie, et al.
Pubblicazione: (2023)
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
di: Chung, Woo-Jin, et al.
Pubblicazione: (2023)
di: Chung, Woo-Jin, et al.
Pubblicazione: (2023)
A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation
di: Taenzer, Michael
Pubblicazione: (2026)
di: Taenzer, Michael
Pubblicazione: (2026)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
di: Sasu, David, et al.
Pubblicazione: (2025)
di: Sasu, David, et al.
Pubblicazione: (2025)
Robust Pitch Estimation and Tracking for Speakers Based on Subband Encoding and the Generalized Labeled Multi-Bernoulli Filter
di: Lin, Shoufeng
Pubblicazione: (2026)
di: Lin, Shoufeng
Pubblicazione: (2026)
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
di: Shirahata, Yuma, et al.
Pubblicazione: (2024)
di: Shirahata, Yuma, et al.
Pubblicazione: (2024)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
di: Wang, Jianzong, et al.
Pubblicazione: (2024)
di: Wang, Jianzong, et al.
Pubblicazione: (2024)
Incremental FastPitch: Chunk-based High Quality Text to Speech
di: Du, Muyang, et al.
Pubblicazione: (2024)
di: Du, Muyang, et al.
Pubblicazione: (2024)
Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis
di: Nippert, Lars
Pubblicazione: (2025)
di: Nippert, Lars
Pubblicazione: (2025)
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
Pitch Estimation With Mean Averaging Smoothed Product Spectrum And Musical Consonance Evaluation Using MASP
di: Baskin, Murat Yasar
Pubblicazione: (2025)
di: Baskin, Murat Yasar
Pubblicazione: (2025)
DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
di: Wei, Haojie, et al.
Pubblicazione: (2024)
di: Wei, Haojie, et al.
Pubblicazione: (2024)
Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception
di: Chen, Jiaxin, et al.
Pubblicazione: (2025)
di: Chen, Jiaxin, et al.
Pubblicazione: (2025)
Periodicity Pitch Detection in Complex Harmonies on EEG Timeline Data
di: Heinze, Maria, et al.
Pubblicazione: (2020)
di: Heinze, Maria, et al.
Pubblicazione: (2020)
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
di: Gu, Yicheng, et al.
Pubblicazione: (2025)
di: Gu, Yicheng, et al.
Pubblicazione: (2025)
Engraving Oriented Joint Estimation of Pitch Spelling and Local and Global Keys
di: Bouquillard, Augustin, et al.
Pubblicazione: (2024)
di: Bouquillard, Augustin, et al.
Pubblicazione: (2024)
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
di: Răgman, Teodora, et al.
Pubblicazione: (2024)
di: Răgman, Teodora, et al.
Pubblicazione: (2024)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
di: Bai, Bingsong, et al.
Pubblicazione: (2024)
di: Bai, Bingsong, et al.
Pubblicazione: (2024)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
di: Kim, Tae-Woo, et al.
Pubblicazione: (2022)
di: Kim, Tae-Woo, et al.
Pubblicazione: (2022)
TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
di: Wang, Yiwen, et al.
Pubblicazione: (2024)
di: Wang, Yiwen, et al.
Pubblicazione: (2024)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
di: Shi, Yu-Fei, et al.
Pubblicazione: (2024)
di: Shi, Yu-Fei, et al.
Pubblicazione: (2024)
PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
di: Riou, Alain, et al.
Pubblicazione: (2025)
di: Riou, Alain, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
di: Ohnaka, Hien, et al.
Pubblicazione: (2026) -
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
di: Kawamura, Masaya, et al.
Pubblicazione: (2025) -
Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity
di: Subramani, Krishna, et al.
Pubblicazione: (2023) -
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024) -
Toward Fully Self-Supervised Multi-Pitch Estimation
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2024)