Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Manabe, Toranosuke, Shibata, Yuto, Takamichi, Shinnosuke, Aoki, Yoshimitsu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
von: Shibata, Yuto, et al.
Veröffentlicht: (2025)
von: Shibata, Yuto, et al.
Veröffentlicht: (2025)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
ProLAP: Probabilistic Language-Audio Pre-Training
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2025)
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2025)
Acoustic-based 3D Human Pose Estimation Robust to Human Position
von: Oumi, Yusuke, et al.
Veröffentlicht: (2024)
von: Oumi, Yusuke, et al.
Veröffentlicht: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds
von: Shibata, Yuto, et al.
Veröffentlicht: (2025)
von: Shibata, Yuto, et al.
Veröffentlicht: (2025)
Brain-to-Speech: Prosody Feature Engineering and Transformer-Based Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2026)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2026)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023)
von: Xin, Detai, et al.
Veröffentlicht: (2023)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
Listening without Looking: Modality Bias in Audio-Visual Captioning
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
YODAS: Youtube-Oriented Dataset for Audio and Speech
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
DNN-based ensemble singing voice synthesis with interactions between singers
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
Spatial-CLAP: Learning Spatially-Aware audio--text Embeddings for Multi-Source Conditions
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
von: Wang, Dingdong, et al.
Veröffentlicht: (2026)
von: Wang, Dingdong, et al.
Veröffentlicht: (2026)
Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification
von: Wu, Bin, et al.
Veröffentlicht: (2024)
von: Wu, Bin, et al.
Veröffentlicht: (2024)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2022)
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2022)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
mmWave Radar Aware Dual-Conditioned GAN for Speech Reconstruction of Signals With Low SNR
von: Karani, Jash, et al.
Veröffentlicht: (2026)
von: Karani, Jash, et al.
Veröffentlicht: (2026)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2
von: Mohanty, Suvendu Sekhar
Veröffentlicht: (2026)
von: Mohanty, Suvendu Sekhar
Veröffentlicht: (2026)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025) -
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025) -
Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
von: Shibata, Yuto, et al.
Veröffentlicht: (2025) -
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025) -
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)