HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Yicheng, Zarazaga, Pablo Pérez, Wang, Chaoren, Wu, Zhizheng, Malisz, Zofia, Henter, Gustav Eje, Juvela, Lauri |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Aliasing-Free Neural Audio Synthesis
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Solid State Bus-Comp: A Large-Scale and Diverse Dataset for Dynamic Range Compressor Virtual Analog Modeling
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
Voice Conversion-based Privacy through Adversarial Information Hiding
von: Webber, Jacob J, et al.
Veröffentlicht: (2024)
von: Webber, Jacob J, et al.
Veröffentlicht: (2024)
HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial Networks
von: Dmitrieva, Ekaterina, et al.
Veröffentlicht: (2025)
von: Dmitrieva, Ekaterina, et al.
Veröffentlicht: (2025)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
De-crackling Virtual Analog Controls with Asymptotically Stable Recurrent Neural Networks
von: Kallinen, Valtteri, et al.
Veröffentlicht: (2025)
von: Kallinen, Valtteri, et al.
Veröffentlicht: (2025)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Collaborative Watermarking for Adversarial Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2023)
von: Juvela, Lauri, et al.
Veröffentlicht: (2023)
Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
von: Han, Changjin, et al.
Veröffentlicht: (2024)
von: Han, Changjin, et al.
Veröffentlicht: (2024)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
Open-Amp: Synthetic Data Framework for Audio Effect Foundation Models
von: Wright, Alec, et al.
Veröffentlicht: (2024)
von: Wright, Alec, et al.
Veröffentlicht: (2024)
Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
Closing the Modality Reasoning Gap for Speech Large Language Models
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
von: Xue, Liumeng, et al.
Veröffentlicht: (2024)
von: Xue, Liumeng, et al.
Veröffentlicht: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
von: He, Haorui, et al.
Veröffentlicht: (2025)
von: He, Haorui, et al.
Veröffentlicht: (2025)
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
von: Guo, Yu-Ren, et al.
Veröffentlicht: (2025)
von: Guo, Yu-Ren, et al.
Veröffentlicht: (2025)
An Initial Investigation of Neural Replay Simulator for Over-the-Air Adversarial Perturbations to Automatic Speaker Verification
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
von: Gu, Yicheng, et al.
Veröffentlicht: (2025) -
Aliasing-Free Neural Audio Synthesis
von: Gu, Yicheng, et al.
Veröffentlicht: (2025) -
Solid State Bus-Comp: A Large-Scale and Diverse Dataset for Dynamic Range Compressor Virtual Analog Modeling
von: Gu, Yicheng, et al.
Veröffentlicht: (2025) -
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025) -
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)