Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nakata, Wataru, Saito, Yuki, Ueda, Yota, Saruwatari, Hiroshi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
Geneses: Unified Generative Speech Enhancement and Separation
von: Asai, Kohei, et al.
Veröffentlicht: (2026)
von: Asai, Kohei, et al.
Veröffentlicht: (2026)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
von: Baba, Kaito, et al.
Veröffentlicht: (2024)
von: Baba, Kaito, et al.
Veröffentlicht: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023)
von: Xin, Detai, et al.
Veröffentlicht: (2023)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
ReverbMiipher: Generative Speech Restoration meets Reverberation Characteristics Controllability
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
Real-time Speech Extraction Using Spatially Regularized Independent Low-rank Matrix Analysis and Rank-constrained Spatial Covariance Matrix Estimation
von: Ishikawa, Yuto, et al.
Veröffentlicht: (2024)
von: Ishikawa, Yuto, et al.
Veröffentlicht: (2024)
Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
von: Wada, Aogu, et al.
Veröffentlicht: (2025)
von: Wada, Aogu, et al.
Veröffentlicht: (2025)
Binaural rendering from microphone array signals of arbitrary geometry
von: Iijima, Naoto, et al.
Veröffentlicht: (2021)
von: Iijima, Naoto, et al.
Veröffentlicht: (2021)
Localizing Acoustic Energy in Sound Field Synthesis by Directionally Weighted Exterior Radiation Suppression
von: Tomita, Yoshihide, et al.
Veröffentlicht: (2024)
von: Tomita, Yoshihide, et al.
Veröffentlicht: (2024)
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
von: Yang, Dong, et al.
Veröffentlicht: (2024)
von: Yang, Dong, et al.
Veröffentlicht: (2024)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2026)
Local Equivariance Error-Based Metrics for Evaluating Sampling-Frequency-Independent Property of Neural Network
von: Imamura, Kanami, et al.
Veröffentlicht: (2025)
von: Imamura, Kanami, et al.
Veröffentlicht: (2025)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
DNN-based ensemble singing voice synthesis with interactions between singers
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides
von: Imamura, Kanami, et al.
Veröffentlicht: (2023)
von: Imamura, Kanami, et al.
Veröffentlicht: (2023)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Construction and Analysis of Impression Caption Dataset for Environmental Sounds
von: Okamoto, Yuki, et al.
Veröffentlicht: (2024)
von: Okamoto, Yuki, et al.
Veröffentlicht: (2024)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
von: Nishigori, Shuichiro, et al.
Veröffentlicht: (2025)
von: Nishigori, Shuichiro, et al.
Veröffentlicht: (2025)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
von: Nakata, Wataru, et al.
Veröffentlicht: (2026) -
Geneses: Unified Generative Speech Enhancement and Separation
von: Asai, Kohei, et al.
Veröffentlicht: (2026) -
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024) -
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
von: Baba, Kaito, et al.
Veröffentlicht: (2024) -
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)