Prosody Analysis of Audiobooks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pethe, Charuta, Pham, Bach, Childress, Felix D, Yin, Yunting, Skiena, Steven |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
Usefulness of Emotional Prosody in Neural Machine Translation
von: Brazier, Charles, et al.
Veröffentlicht: (2024)
von: Brazier, Charles, et al.
Veröffentlicht: (2024)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
PBSCR: The Piano Bootleg Score Composer Recognition Dataset
von: Jain, Arhan, et al.
Veröffentlicht: (2024)
von: Jain, Arhan, et al.
Veröffentlicht: (2024)
Safeguarding Privacy in Edge Speech Understanding with Tiny Foundation Models
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
von: Benazir, Afsara, et al.
Veröffentlicht: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
von: Park, Kyeongman, et al.
Veröffentlicht: (2025)
von: Park, Kyeongman, et al.
Veröffentlicht: (2025)
Prosody-Driven Privacy-Preserving Dementia Detection
von: Woszczyk, Dominika, et al.
Veröffentlicht: (2024)
von: Woszczyk, Dominika, et al.
Veröffentlicht: (2024)
Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
von: Haider, Daniel, et al.
Veröffentlicht: (2024)
von: Haider, Daniel, et al.
Veröffentlicht: (2024)
Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech
von: Agbo, Idoko, et al.
Veröffentlicht: (2024)
von: Agbo, Idoko, et al.
Veröffentlicht: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
von: Thinakaran, Preethi, et al.
Veröffentlicht: (2025)
von: Thinakaran, Preethi, et al.
Veröffentlicht: (2025)
VANPY: Voice Analysis Framework
von: Koushnir, Gregory, et al.
Veröffentlicht: (2025)
von: Koushnir, Gregory, et al.
Veröffentlicht: (2025)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
von: Yin, Chun, et al.
Veröffentlicht: (2024)
von: Yin, Chun, et al.
Veröffentlicht: (2024)
An Analysis of the Variance of Diffusion-based Speech Enhancement
von: Lay, Bunlong, et al.
Veröffentlicht: (2024)
von: Lay, Bunlong, et al.
Veröffentlicht: (2024)
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
von: Sohn, Samuel S., et al.
Veröffentlicht: (2025)
von: Sohn, Samuel S., et al.
Veröffentlicht: (2025)
MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
Cross-domain Sound Recognition for Efficient Underwater Data Analysis
von: Park, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Park, Jeongsoo, et al.
Veröffentlicht: (2023)
Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2024)
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2024)
A Data-Driven Analysis of Robust Automatic Piano Transcription
von: Edwards, Drew, et al.
Veröffentlicht: (2024)
von: Edwards, Drew, et al.
Veröffentlicht: (2024)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2024)
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
A Comprehensive Survey on Heart Sound Analysis in the Deep Learning Era
von: Ren, Zhao, et al.
Veröffentlicht: (2023)
von: Ren, Zhao, et al.
Veröffentlicht: (2023)
I Guess That's Why They Call it the Blues: Causal Analysis for Audio Classifiers
von: Kelly, David A., et al.
Veröffentlicht: (2026)
von: Kelly, David A., et al.
Veröffentlicht: (2026)
Analysis-Driven Procedural Generation of an Engine Sound Dataset with Embedded Control Annotations
von: Doerfler, Robin, et al.
Veröffentlicht: (2026)
von: Doerfler, Robin, et al.
Veröffentlicht: (2026)
StethoLM: Audio Language Model for Cardiopulmonary Analysis Across Clinical Tasks
von: Wang, Yishan, et al.
Veröffentlicht: (2026)
von: Wang, Yishan, et al.
Veröffentlicht: (2026)
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
von: Richter, Julius, et al.
Veröffentlicht: (2024)
von: Richter, Julius, et al.
Veröffentlicht: (2024)
ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis
von: Ni-Hahn, Stephen, et al.
Veröffentlicht: (2025)
von: Ni-Hahn, Stephen, et al.
Veröffentlicht: (2025)
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
Modeling Analog Dynamic Range Compressors using Deep Learning and State-space Models
von: Yin, Hanzhi, et al.
Veröffentlicht: (2024)
von: Yin, Hanzhi, et al.
Veröffentlicht: (2024)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
An Ensemble Approach to Music Source Separation: A Comparative Analysis of Conventional and Hierarchical Stem Separation
von: Vardhan, Saarth, et al.
Veröffentlicht: (2024)
von: Vardhan, Saarth, et al.
Veröffentlicht: (2024)
Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech
von: Sun, Anchen, et al.
Veröffentlicht: (2025)
von: Sun, Anchen, et al.
Veröffentlicht: (2025)
Detection of Electric Motor Damage Through Analysis of Sound Signals Using Bayesian Neural Networks
von: Bauer, Waldemar, et al.
Veröffentlicht: (2024)
von: Bauer, Waldemar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026) -
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025) -
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025) -
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025) -
Usefulness of Emotional Prosody in Neural Machine Translation
von: Brazier, Charles, et al.
Veröffentlicht: (2024)