PitchFlower: A flow-based neural audio codec with pitch controllability
Fuente:
arXiv
Salvato in:
| Autori principali: | Torres, Diego, Roebel, Axel, Obin, Nicolas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
di: Abrassart, Mathilde, et al.
Pubblicazione: (2025)
di: Abrassart, Mathilde, et al.
Pubblicazione: (2025)
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
di: Pepino, Leonardo, et al.
Pubblicazione: (2023)
di: Pepino, Leonardo, et al.
Pubblicazione: (2023)
FlowDec: A flow-based full-band general audio codec with high perceptual quality
di: Welker, Simon, et al.
Pubblicazione: (2025)
di: Welker, Simon, et al.
Pubblicazione: (2025)
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
LDCodec: A high quality neural audio codec with low-complexity decoder
di: Jiang, Jiawei, et al.
Pubblicazione: (2025)
di: Jiang, Jiawei, et al.
Pubblicazione: (2025)
Speaker anonymization using neural audio codec language models
di: Panariello, Michele, et al.
Pubblicazione: (2023)
di: Panariello, Michele, et al.
Pubblicazione: (2023)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model
di: Hauret, Julien, et al.
Pubblicazione: (2025)
di: Hauret, Julien, et al.
Pubblicazione: (2025)
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
di: Siuzdak, Hubert
Pubblicazione: (2023)
di: Siuzdak, Hubert
Pubblicazione: (2023)
Combining audio control and style transfer using latent diffusion
di: Demerlé, Nils, et al.
Pubblicazione: (2024)
di: Demerlé, Nils, et al.
Pubblicazione: (2024)
Robustness of Speech Separation Models for Similar-pitch Speakers
di: Lay, Bunlong, et al.
Pubblicazione: (2024)
di: Lay, Bunlong, et al.
Pubblicazione: (2024)
Transformation of audio embeddings into interpretable, concept-based representations
di: Zhang, Alice, et al.
Pubblicazione: (2025)
di: Zhang, Alice, et al.
Pubblicazione: (2025)
Room-acoustic simulations as an alternative to measurements for audio-algorithm evaluation
di: Götz, Georg, et al.
Pubblicazione: (2025)
di: Götz, Georg, et al.
Pubblicazione: (2025)
Incremental learning for audio classification with Hebbian Deep Neural Networks
di: Casciotti, Riccardo, et al.
Pubblicazione: (2026)
di: Casciotti, Riccardo, et al.
Pubblicazione: (2026)
Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4
di: Park, Jongyeon, et al.
Pubblicazione: (2025)
di: Park, Jongyeon, et al.
Pubblicazione: (2025)
Toward Fully Self-Supervised Multi-Pitch Estimation
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2024)
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2024)
FreeCodec: A disentangled neural speech codec with fewer tokens
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2025)
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2025)
Towards generalizing deep-audio fake detection networks
di: Gasenzer, Konstantin, et al.
Pubblicazione: (2023)
di: Gasenzer, Konstantin, et al.
Pubblicazione: (2023)
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2025)
di: Cwitkowitz, Frank, et al.
Pubblicazione: (2025)
HyperGANStrument: Instrument Sound Synthesis and Editing with Pitch-Invariant Hypernetworks
di: Zhang, Zhe, et al.
Pubblicazione: (2024)
di: Zhang, Zhe, et al.
Pubblicazione: (2024)
Testing chatbots on the creation of encoders for audio conditioned image generation
di: León, Jorge E., et al.
Pubblicazione: (2025)
di: León, Jorge E., et al.
Pubblicazione: (2025)
Unsupervised outlier detection to improve bird audio dataset labels
di: Collins, Bruce
Pubblicazione: (2025)
di: Collins, Bruce
Pubblicazione: (2025)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
di: Rouard, Simon, et al.
Pubblicazione: (2025)
di: Rouard, Simon, et al.
Pubblicazione: (2025)
Late fusion ensembles for speech recognition on diverse input audio representations
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
di: Jeziorek, Kamil, et al.
Pubblicazione: (2026)
di: Jeziorek, Kamil, et al.
Pubblicazione: (2026)
Versatile audio-visual learning for emotion recognition
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
Continuous Audio Language Models
di: Rouard, Simon, et al.
Pubblicazione: (2025)
di: Rouard, Simon, et al.
Pubblicazione: (2025)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
di: Rouard, Simon, et al.
Pubblicazione: (2024)
di: Rouard, Simon, et al.
Pubblicazione: (2024)
Sentiment analysis in non-fixed length audios using a Fully Convolutional Neural Network
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
Decodable but not structured: linear probing enables Underwater Acoustic Target Recognition with pretrained audio embeddings
di: Hummel, Hilde I., et al.
Pubblicazione: (2026)
di: Hummel, Hilde I., et al.
Pubblicazione: (2026)
Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
di: Lin, Tsung-En, et al.
Pubblicazione: (2025)
di: Lin, Tsung-En, et al.
Pubblicazione: (2025)
PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
di: Riou, Alain, et al.
Pubblicazione: (2025)
di: Riou, Alain, et al.
Pubblicazione: (2025)
Spectrogram features for audio and speech analysis
di: McLoughlin, Ian, et al.
Pubblicazione: (2026)
di: McLoughlin, Ian, et al.
Pubblicazione: (2026)
Learning Relationships Between Separate Audio Tracks for Creative Applications
di: Bujard, Balthazar, et al.
Pubblicazione: (2025)
di: Bujard, Balthazar, et al.
Pubblicazione: (2025)
PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
di: Hono, Yukiya, et al.
Pubblicazione: (2024)
di: Hono, Yukiya, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
di: Abrassart, Mathilde, et al.
Pubblicazione: (2025) -
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
di: Lemerle, Théodor, et al.
Pubblicazione: (2024) -
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
di: Pepino, Leonardo, et al.
Pubblicazione: (2023) -
FlowDec: A flow-based full-band general audio codec with high perceptual quality
di: Welker, Simon, et al.
Pubblicazione: (2025) -
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)