A Context-Based Numerical Format Prediction for a Text-To-Speech System
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Darwesh, Yaser, Wern, Lit Wei, Mustafa, Mumtaz Begum |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
von: Richter, Julius, et al.
Veröffentlicht: (2026)
von: Richter, Julius, et al.
Veröffentlicht: (2026)
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
von: Pešán, Jan, et al.
Veröffentlicht: (2024)
von: Pešán, Jan, et al.
Veröffentlicht: (2024)
ASTRA: Aligning Speech and Text Representations for Asr without Sampling
von: Gaur, Neeraj, et al.
Veröffentlicht: (2024)
von: Gaur, Neeraj, et al.
Veröffentlicht: (2024)
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
von: Chen, Zehua, et al.
Veröffentlicht: (2023)
von: Chen, Zehua, et al.
Veröffentlicht: (2023)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2026)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2026)
Latent-Domain Predictive Neural Speech Coding
von: Jiang, Xue, et al.
Veröffentlicht: (2022)
von: Jiang, Xue, et al.
Veröffentlicht: (2022)
Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
von: Khorrami, Khazar, et al.
Veröffentlicht: (2023)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2023)
Text to Speech System for Meitei Mayek Script
von: Irengbam, Gangular Singh, et al.
Veröffentlicht: (2025)
von: Irengbam, Gangular Singh, et al.
Veröffentlicht: (2025)
Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior
von: Yemini, Yochai, et al.
Veröffentlicht: (2025)
von: Yemini, Yochai, et al.
Veröffentlicht: (2025)
Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
von: Gerazov, Branislav, et al.
Veröffentlicht: (2025)
von: Gerazov, Branislav, et al.
Veröffentlicht: (2025)
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
von: Lovelace, Justin, et al.
Veröffentlicht: (2025)
von: Lovelace, Justin, et al.
Veröffentlicht: (2025)
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
von: Cao, Fengyuan, et al.
Veröffentlicht: (2026)
von: Cao, Fengyuan, et al.
Veröffentlicht: (2026)
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus with Punctuation, Casing, and Context
von: Povey, Anna, et al.
Veröffentlicht: (2024)
von: Povey, Anna, et al.
Veröffentlicht: (2024)
Speech Understanding on Tiny Devices with A Learning Cache
von: Benazir, Afsara, et al.
Veröffentlicht: (2023)
von: Benazir, Afsara, et al.
Veröffentlicht: (2023)
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?
von: Makarov, Rostislav, et al.
Veröffentlicht: (2025)
von: Makarov, Rostislav, et al.
Veröffentlicht: (2025)
Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
Vibravox: A Dataset of French Speech Captured with Body-conduction Audio Sensors
von: Hauret, Julien, et al.
Veröffentlicht: (2024)
von: Hauret, Julien, et al.
Veröffentlicht: (2024)
Discrete-Time Diffusion-Like Models for Speech Synthesis
von: Tan, Xiaozhou, et al.
Veröffentlicht: (2025)
von: Tan, Xiaozhou, et al.
Veröffentlicht: (2025)
Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
von: Yanuka, Moran, et al.
Veröffentlicht: (2025)
von: Yanuka, Moran, et al.
Veröffentlicht: (2025)
Toward a Reinforcement-Learning-Based System for Adjusting Medication to Minimize Speech Disfluency
von: Constas, Pavlos, et al.
Veröffentlicht: (2023)
von: Constas, Pavlos, et al.
Veröffentlicht: (2023)
RepCodec: A Speech Representation Codec for Speech Tokenization
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
von: Close, George, et al.
Veröffentlicht: (2025)
von: Close, George, et al.
Veröffentlicht: (2025)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
Coverage-Guaranteed Speech Emotion Recognition via Calibrated Uncertainty-Adaptive Prediction Sets
von: Jia, Zijun, et al.
Veröffentlicht: (2025)
von: Jia, Zijun, et al.
Veröffentlicht: (2025)
Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
von: Sadeghi, Mostafa, et al.
Veröffentlicht: (2025)
von: Sadeghi, Mostafa, et al.
Veröffentlicht: (2025)
Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture
von: Ouyang, Qianhe
Veröffentlicht: (2025)
von: Ouyang, Qianhe
Veröffentlicht: (2025)
Ähnliche Einträge
-
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
von: Glazer, Neta, et al.
Veröffentlicht: (2025) -
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023) -
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024) -
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025) -
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
von: Richter, Julius, et al.
Veröffentlicht: (2026)