SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saito, Yuki, Igarashi, Takuto, Seki, Kentaro, Takamichi, Shinnosuke, Yamamoto, Ryuichi, Tachibana, Kentaro, Saruwatari, Hiroshi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023)
von: Xin, Detai, et al.
Veröffentlicht: (2023)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
von: Take, Osamu, et al.
Veröffentlicht: (2024)
von: Take, Osamu, et al.
Veröffentlicht: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2022)
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2022)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
DNN-based ensemble singing voice synthesis with interactions between singers
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
von: Kawamura, Masaya, et al.
Veröffentlicht: (2024)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2024)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
Geneses: Unified Generative Speech Enhancement and Separation
von: Asai, Kohei, et al.
Veröffentlicht: (2026)
von: Asai, Kohei, et al.
Veröffentlicht: (2026)
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
von: Baba, Kaito, et al.
Veröffentlicht: (2024)
von: Baba, Kaito, et al.
Veröffentlicht: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
Binaural rendering from microphone array signals of arbitrary geometry
von: Iijima, Naoto, et al.
Veröffentlicht: (2021)
von: Iijima, Naoto, et al.
Veröffentlicht: (2021)
Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
von: Wada, Aogu, et al.
Veröffentlicht: (2025)
von: Wada, Aogu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024) -
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024) -
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024) -
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025) -
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023)