Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Yicheng, Wang, Chaoren, Wu, Zhizheng, Juvela, Lauri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aliasing-Free Neural Audio Synthesis
by: Gu, Yicheng, et al.
Published: (2025)
by: Gu, Yicheng, et al.
Published: (2025)
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
by: Gu, Yicheng, et al.
Published: (2024)
by: Gu, Yicheng, et al.
Published: (2024)
Solid State Bus-Comp: A Large-Scale and Diverse Dataset for Dynamic Range Compressor Virtual Analog Modeling
by: Gu, Yicheng, et al.
Published: (2025)
by: Gu, Yicheng, et al.
Published: (2025)
De-crackling Virtual Analog Controls with Asymptotically Stable Recurrent Neural Networks
by: Kallinen, Valtteri, et al.
Published: (2025)
by: Kallinen, Valtteri, et al.
Published: (2025)
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
by: Gu, Yicheng, et al.
Published: (2025)
by: Gu, Yicheng, et al.
Published: (2025)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
by: Juvela, Lauri, et al.
Published: (2024)
by: Juvela, Lauri, et al.
Published: (2024)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
by: Gu, Yicheng, et al.
Published: (2024)
by: Gu, Yicheng, et al.
Published: (2024)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
by: Tee, Hitomi Jin Ling, et al.
Published: (2025)
by: Tee, Hitomi Jin Ling, et al.
Published: (2025)
Collaborative Watermarking for Adversarial Speech Synthesis
by: Juvela, Lauri, et al.
Published: (2023)
by: Juvela, Lauri, et al.
Published: (2023)
Improving Neural Pitch Estimation with SWIPE Kernels
by: Marttila, David, et al.
Published: (2025)
by: Marttila, David, et al.
Published: (2025)
Cross-domain Neural Pitch and Periodicity Estimation
by: Morrison, Max, et al.
Published: (2023)
by: Morrison, Max, et al.
Published: (2023)
Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
by: Zhang, Xueyao, et al.
Published: (2023)
by: Zhang, Xueyao, et al.
Published: (2023)
An Initial Investigation of Neural Replay Simulator for Over-the-Air Adversarial Perturbations to Automatic Speaker Verification
by: Li, Jiaqi, et al.
Published: (2023)
by: Li, Jiaqi, et al.
Published: (2023)
Closing the Modality Reasoning Gap for Speech Large Language Models
by: Wang, Chaoren, et al.
Published: (2026)
by: Wang, Chaoren, et al.
Published: (2026)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
by: Xue, Liumeng, et al.
Published: (2024)
by: Xue, Liumeng, et al.
Published: (2024)
End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models
by: Juvela, Lauri, et al.
Published: (2024)
by: Juvela, Lauri, et al.
Published: (2024)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
by: Zhang, Xueyao, et al.
Published: (2023)
by: Zhang, Xueyao, et al.
Published: (2023)
TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
by: Wang, Yiwen, et al.
Published: (2024)
by: Wang, Yiwen, et al.
Published: (2024)
CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation
by: Liang, Ziqi, et al.
Published: (2025)
by: Liang, Ziqi, et al.
Published: (2025)
Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity
by: Subramani, Krishna, et al.
Published: (2023)
by: Subramani, Krishna, et al.
Published: (2023)
Open-Amp: Synthetic Data Framework for Audio Effect Foundation Models
by: Wright, Alec, et al.
Published: (2024)
by: Wright, Alec, et al.
Published: (2024)
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
by: He, Haorui, et al.
Published: (2025)
by: He, Haorui, et al.
Published: (2025)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
by: Wang, Jianzong, et al.
Published: (2024)
by: Wang, Jianzong, et al.
Published: (2024)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
by: Kim, Tae-Woo, et al.
Published: (2022)
by: Kim, Tae-Woo, et al.
Published: (2022)
Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement
by: Li, Yi, et al.
Published: (2024)
by: Li, Yi, et al.
Published: (2024)
Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception
by: Chen, Jiaxin, et al.
Published: (2025)
by: Chen, Jiaxin, et al.
Published: (2025)
Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach
by: Abeßer, Jakob, et al.
Published: (2025)
by: Abeßer, Jakob, et al.
Published: (2025)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
by: Riou, Alain, et al.
Published: (2023)
by: Riou, Alain, et al.
Published: (2023)
Periodicity Pitch Detection in Complex Harmonies on EEG Timeline Data
by: Heinze, Maria, et al.
Published: (2020)
by: Heinze, Maria, et al.
Published: (2020)
Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
by: Yang, Peiji, et al.
Published: (2024)
by: Yang, Peiji, et al.
Published: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
by: Zhou, Xuehao, et al.
Published: (2024)
by: Zhou, Xuehao, et al.
Published: (2024)
Harmonic Summation-Based Robust Pitch Estimation in Noisy and Reverberant Environments
by: Singh, Anup, et al.
Published: (2025)
by: Singh, Anup, et al.
Published: (2025)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
by: Wei, Haojie, et al.
Published: (2023)
by: Wei, Haojie, et al.
Published: (2023)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
by: Shi, Yu-Fei, et al.
Published: (2024)
by: Shi, Yu-Fei, et al.
Published: (2024)
A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation
by: Taenzer, Michael
Published: (2026)
by: Taenzer, Michael
Published: (2026)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
by: Wagner, Dominik, et al.
Published: (2023)
by: Wagner, Dominik, et al.
Published: (2023)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
by: Chen, Meiying, et al.
Published: (2022)
by: Chen, Meiying, et al.
Published: (2022)
Similar Items
-
Aliasing-Free Neural Audio Synthesis
by: Gu, Yicheng, et al.
Published: (2025) -
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
by: Gu, Yicheng, et al.
Published: (2024) -
Solid State Bus-Comp: A Large-Scale and Diverse Dataset for Dynamic Range Compressor Virtual Analog Modeling
by: Gu, Yicheng, et al.
Published: (2025) -
De-crackling Virtual Analog Controls with Asymptotically Stable Recurrent Neural Networks
by: Kallinen, Valtteri, et al.
Published: (2025) -
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
by: Gu, Yicheng, et al.
Published: (2025)