An Attribute Interpolation Method in Speech Synthesis by Model Merging
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Murata, Masato, Miyazaki, Koichi, Koriyama, Tomoki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Eigenvoice Synthesis based on Model Editing for Speaker Generation
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
Exploring the Capability of Mamba in Speech Applications
von: Miyazaki, Koichi, et al.
Veröffentlicht: (2024)
von: Miyazaki, Koichi, et al.
Veröffentlicht: (2024)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
von: Koriyama, Tomoki
Veröffentlicht: (2024)
von: Koriyama, Tomoki
Veröffentlicht: (2024)
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
von: Yang, Dong, et al.
Veröffentlicht: (2024)
von: Yang, Dong, et al.
Veröffentlicht: (2024)
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
von: Nishigori, Shuichiro, et al.
Veröffentlicht: (2025)
von: Nishigori, Shuichiro, et al.
Veröffentlicht: (2025)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
Diffusion-based Signal Refiner for Speech Enhancement and Separation
von: Hirano, Masato, et al.
Veröffentlicht: (2023)
von: Hirano, Masato, et al.
Veröffentlicht: (2023)
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
Electrolaryngeal Speech Intelligibility Enhancement Through Robust Linguistic Encoders
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2023)
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2023)
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
Wavehax: Aliasing-Free Neural Waveform Synthesis Based on 2D Convolution and Harmonic Prior for Reliable Complex Spectrogram Estimation
von: Yoneyama, Reo, et al.
Veröffentlicht: (2024)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2024)
Serial-OE: Anomalous sound detection based on serial method with outlier exposure capable of using small amounts of anomalous data for training
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Eigenvoice Synthesis based on Model Editing for Speaker Generation
von: Murata, Masato, et al.
Veröffentlicht: (2025) -
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025) -
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024) -
Exploring the Capability of Mamba in Speech Applications
von: Miyazaki, Koichi, et al.
Veröffentlicht: (2024) -
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)