Analysis of the Maximum Prediction Gain of Short-Term Prediction on Sustained Speech
Fuente:
arXiv
Guardado en:
| Autores principales: | Hinrichs, Reemt, Damara, Muhamad Fadli, Preihs, Stephan, Ostermann, Jörn |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pruning-aware Loss Functions for STOI-Optimized Pruned Recurrent Autoencoders for the Compression of the Stimulation Patterns of Cochlear Implants at Zero Delay
por: Hinrichs, Reemt, et al.
Publicado: (2025)
por: Hinrichs, Reemt, et al.
Publicado: (2025)
A Dataset for Automatic Vocal Mode Classification
por: Hinrichs, Reemt, et al.
Publicado: (2026)
por: Hinrichs, Reemt, et al.
Publicado: (2026)
LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
por: Jain, Arnav, et al.
Publicado: (2024)
por: Jain, Arnav, et al.
Publicado: (2024)
Future Full-Ocean Deep SSPs Prediction based on Hierarchical Long Short-Term Memory Neural Networks
por: Lu, Jiajun, et al.
Publicado: (2023)
por: Lu, Jiajun, et al.
Publicado: (2023)
Perceived Femininity in Singing Voice: Analysis and Prediction
por: Kong, Yuexuan, et al.
Publicado: (2025)
por: Kong, Yuexuan, et al.
Publicado: (2025)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Combined Generative and Predictive Modeling for Speech Super-resolution
por: Wang, Heming, et al.
Publicado: (2024)
por: Wang, Heming, et al.
Publicado: (2024)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
por: Zhao, Yiyang, et al.
Publicado: (2025)
por: Zhao, Yiyang, et al.
Publicado: (2025)
Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2024)
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2024)
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
por: Engert, Natalie, et al.
Publicado: (2026)
por: Engert, Natalie, et al.
Publicado: (2026)
DeepGESI: A Non-Intrusive Objective Evaluation Model for Predicting Speech Intelligibility in Hearing-Impaired Listeners
por: Luo, Wenyu, et al.
Publicado: (2025)
por: Luo, Wenyu, et al.
Publicado: (2025)
Harmonic Detection from Noisy Speech with Auditory Frame Gain for Intelligibility Enhancement
por: Queiroz, A., et al.
Publicado: (2024)
por: Queiroz, A., et al.
Publicado: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
por: Zhang, Leying, et al.
Publicado: (2025)
por: Zhang, Leying, et al.
Publicado: (2025)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
por: Huang, Wen-Chin, et al.
Publicado: (2024)
por: Huang, Wen-Chin, et al.
Publicado: (2024)
Faster Speech-LLaMA Inference with Multi-token Prediction
por: Raj, Desh, et al.
Publicado: (2024)
por: Raj, Desh, et al.
Publicado: (2024)
Non-Invasive Suicide Risk Prediction Through Speech Analysis
por: Amiriparian, Shahin, et al.
Publicado: (2024)
por: Amiriparian, Shahin, et al.
Publicado: (2024)
A Multi-decoder Neural Tracking Method for Accurately Predicting Speech Intelligibility
por: Sonck, Rien, et al.
Publicado: (2026)
por: Sonck, Rien, et al.
Publicado: (2026)
Latent-Domain Predictive Neural Speech Coding
por: Jiang, Xue, et al.
Publicado: (2022)
por: Jiang, Xue, et al.
Publicado: (2022)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
por: Chen, Li-Wei, et al.
Publicado: (2024)
por: Chen, Li-Wei, et al.
Publicado: (2024)
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
por: Rong, Xiaobin, et al.
Publicado: (2026)
por: Rong, Xiaobin, et al.
Publicado: (2026)
Multi-Utterance Speech Separation and Association Trained on Short Segments
por: Wang, Yuzhu, et al.
Publicado: (2025)
por: Wang, Yuzhu, et al.
Publicado: (2025)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2025)
por: Inoue, Sho, et al.
Publicado: (2025)
Representing Speech Through Autoregressive Prediction of Cochlear Tokens
por: Tuckute, Greta, et al.
Publicado: (2025)
por: Tuckute, Greta, et al.
Publicado: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
por: Wang, Jianjin, et al.
Publicado: (2025)
por: Wang, Jianjin, et al.
Publicado: (2025)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
por: Ai, Yang, et al.
Publicado: (2024)
por: Ai, Yang, et al.
Publicado: (2024)
Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
por: Xue, Ke, et al.
Publicado: (2026)
por: Xue, Ke, et al.
Publicado: (2026)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
por: Yamamoto, Katsuhiko, et al.
Publicado: (2025)
por: Yamamoto, Katsuhiko, et al.
Publicado: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
por: Xu, Mohan, et al.
Publicado: (2024)
por: Xu, Mohan, et al.
Publicado: (2024)
KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
por: Xia, Kangxiang, et al.
Publicado: (2024)
por: Xia, Kangxiang, et al.
Publicado: (2024)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
por: Liu, Fei, et al.
Publicado: (2024)
por: Liu, Fei, et al.
Publicado: (2024)
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
por: Shi, Hao, et al.
Publicado: (2023)
por: Shi, Hao, et al.
Publicado: (2023)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
por: Lin, Zhennan, et al.
Publicado: (2025)
por: Lin, Zhennan, et al.
Publicado: (2025)
No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
por: Zhou, Haoshuai, et al.
Publicado: (2025)
por: Zhou, Haoshuai, et al.
Publicado: (2025)
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
por: Liang, Xinyu, et al.
Publicado: (2025)
por: Liang, Xinyu, et al.
Publicado: (2025)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
por: Zezario, Ryandhimas E., et al.
Publicado: (2025)
por: Zezario, Ryandhimas E., et al.
Publicado: (2025)
Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
por: Qin, Chengyuan, et al.
Publicado: (2025)
por: Qin, Chengyuan, et al.
Publicado: (2025)
Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
por: Zhou, Haoshuai, et al.
Publicado: (2025)
por: Zhou, Haoshuai, et al.
Publicado: (2025)
Emotion-Aware Quantization for Discrete Speech Representations: An Analysis of Emotion Preservation
por: Zhou, Haoguang, et al.
Publicado: (2026)
por: Zhou, Haoguang, et al.
Publicado: (2026)
Articulatory Feature Prediction from Surface EMG during Speech Production
por: Lee, Jihwan, et al.
Publicado: (2025)
por: Lee, Jihwan, et al.
Publicado: (2025)
Ejemplares similares
-
Pruning-aware Loss Functions for STOI-Optimized Pruned Recurrent Autoencoders for the Compression of the Stimulation Patterns of Cochlear Implants at Zero Delay
por: Hinrichs, Reemt, et al.
Publicado: (2025) -
A Dataset for Automatic Vocal Mode Classification
por: Hinrichs, Reemt, et al.
Publicado: (2026) -
LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
por: Jain, Arnav, et al.
Publicado: (2024) -
Future Full-Ocean Deep SSPs Prediction based on Hierarchical Long Short-Term Memory Neural Networks
por: Lu, Jiajun, et al.
Publicado: (2023) -
Perceived Femininity in Singing Voice: Analysis and Prediction
por: Kong, Yuexuan, et al.
Publicado: (2025)