Prosody Labeling with Phoneme-BERT and Speech Foundation Models
Fuente:
arXiv
Guardado en:
| Autor principal: | Koriyama, Tomoki |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
por: Koriyama, Tomoki
Publicado: (2024)
por: Koriyama, Tomoki
Publicado: (2024)
An Attribute Interpolation Method in Speech Synthesis by Model Merging
por: Murata, Masato, et al.
Publicado: (2024)
por: Murata, Masato, et al.
Publicado: (2024)
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
por: Yang, Dong, et al.
Publicado: (2024)
por: Yang, Dong, et al.
Publicado: (2024)
Eigenvoice Synthesis based on Model Editing for Speaker Generation
por: Murata, Masato, et al.
Publicado: (2025)
por: Murata, Masato, et al.
Publicado: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
por: Murata, Masato, et al.
Publicado: (2025)
por: Murata, Masato, et al.
Publicado: (2025)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
por: Karapiperis, Sotirios, et al.
Publicado: (2024)
por: Karapiperis, Sotirios, et al.
Publicado: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
por: Jiang, Yuepeng, et al.
Publicado: (2024)
por: Jiang, Yuepeng, et al.
Publicado: (2024)
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
por: Salvi, Davide, et al.
Publicado: (2025)
por: Salvi, Davide, et al.
Publicado: (2025)
A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
por: Monir, Nasser-Eddine, et al.
Publicado: (2024)
por: Monir, Nasser-Eddine, et al.
Publicado: (2024)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
por: Du, Chenpeng, et al.
Publicado: (2021)
por: Du, Chenpeng, et al.
Publicado: (2021)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
por: Monir, Nasser-Eddine, et al.
Publicado: (2025)
por: Monir, Nasser-Eddine, et al.
Publicado: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
Benchmarking Prosody Encoding in Discrete Speech Tokens
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
por: Huang, Wen-Chin, et al.
Publicado: (2024)
por: Huang, Wen-Chin, et al.
Publicado: (2024)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
por: Zhou, Xuanru, et al.
Publicado: (2025)
por: Zhou, Xuanru, et al.
Publicado: (2025)
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
por: Pan, Jianan, et al.
Publicado: (2026)
por: Pan, Jianan, et al.
Publicado: (2026)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
por: Honda, Tomoki, et al.
Publicado: (2024)
por: Honda, Tomoki, et al.
Publicado: (2024)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
por: Hu, Cheng-Hung, et al.
Publicado: (2025)
por: Hu, Cheng-Hung, et al.
Publicado: (2025)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
por: Borodin, Kirill, et al.
Publicado: (2026)
por: Borodin, Kirill, et al.
Publicado: (2026)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
por: Sutherland, Robert, et al.
Publicado: (2024)
por: Sutherland, Robert, et al.
Publicado: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
por: Ku, Pin-Jui, et al.
Publicado: (2024)
por: Ku, Pin-Jui, et al.
Publicado: (2024)
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
por: Yang, Dong, et al.
Publicado: (2025)
por: Yang, Dong, et al.
Publicado: (2025)
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
por: Medin, Lucas Block, et al.
Publicado: (2025)
por: Medin, Lucas Block, et al.
Publicado: (2025)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
por: Wu, Wenxuan, et al.
Publicado: (2024)
por: Wu, Wenxuan, et al.
Publicado: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
por: Huang, Wen-Chin, et al.
Publicado: (2025)
por: Huang, Wen-Chin, et al.
Publicado: (2025)
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
por: Thinakaran, Preethi, et al.
Publicado: (2025)
por: Thinakaran, Preethi, et al.
Publicado: (2025)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
por: Borodin, Kirill, et al.
Publicado: (2025)
por: Borodin, Kirill, et al.
Publicado: (2025)
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
por: Cooper, Erica, et al.
Publicado: (2025)
por: Cooper, Erica, et al.
Publicado: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
por: He, Xiangheng, et al.
Publicado: (2024)
por: He, Xiangheng, et al.
Publicado: (2024)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
por: Ahn, Hyebin, et al.
Publicado: (2025)
por: Ahn, Hyebin, et al.
Publicado: (2025)
Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
por: Kolani, Yakov, et al.
Publicado: (2025)
por: Kolani, Yakov, et al.
Publicado: (2025)
Multitask Learning for Grapheme-to-Phoneme Conversion of Anglicisms in German Speech Recognition
por: Pritzen, Julia, et al.
Publicado: (2021)
por: Pritzen, Julia, et al.
Publicado: (2021)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
por: Ogg, Mattson, et al.
Publicado: (2025)
por: Ogg, Mattson, et al.
Publicado: (2025)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
por: Chen, Xueyuan, et al.
Publicado: (2024)
por: Chen, Xueyuan, et al.
Publicado: (2024)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
por: Shi, Jiatong, et al.
Publicado: (2023)
por: Shi, Jiatong, et al.
Publicado: (2023)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
por: Ma, Ding, et al.
Publicado: (2026)
por: Ma, Ding, et al.
Publicado: (2026)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
por: Chen, Zhengyang, et al.
Publicado: (2024)
por: Chen, Zhengyang, et al.
Publicado: (2024)
Ejemplares similares
-
VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
por: Koriyama, Tomoki
Publicado: (2024) -
An Attribute Interpolation Method in Speech Synthesis by Model Merging
por: Murata, Masato, et al.
Publicado: (2024) -
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
por: Yang, Dong, et al.
Publicado: (2024) -
Eigenvoice Synthesis based on Model Editing for Speaker Generation
por: Murata, Masato, et al.
Publicado: (2025) -
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
por: Murata, Masato, et al.
Publicado: (2025)