Proposal of protocols for speech materials acquisition and presentation assisted by tools based on structured test signals
Fuente:
arXiv
Guardado en:
| Autores principales: | Kawahara, Hideki, Sakakibara, Ken-Ichi, Mizumachi, Mitsunori, Yatabe, Kohei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sound Safeguarding for Acoustic Measurement Using Any Sounds: Tools and Applications
por: Kawahara, Hideki, et al.
Publicado: (2025)
por: Kawahara, Hideki, et al.
Publicado: (2025)
Subband Splitting: Simple, Efficient and Effective Technique for Solving Block Permutation Problem in Determined Blind Source Separation
por: Matsumoto, Kazuki, et al.
Publicado: (2024)
por: Matsumoto, Kazuki, et al.
Publicado: (2024)
SELEBI: Percussion-aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform
por: Akaishi, Natsuki, et al.
Publicado: (2026)
por: Akaishi, Natsuki, et al.
Publicado: (2026)
Local Equivariance Error-Based Metrics for Evaluating Sampling-Frequency-Independent Property of Neural Network
por: Imamura, Kanami, et al.
Publicado: (2025)
por: Imamura, Kanami, et al.
Publicado: (2025)
Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides
por: Imamura, Kanami, et al.
Publicado: (2023)
por: Imamura, Kanami, et al.
Publicado: (2023)
Interactive tools for making temporally variable, multiple-attributes, and multiple-instances morphing accessible: Flexible manipulation of divergent speech instances for explorational research and education
por: Kawahara, Hideki, et al.
Publicado: (2024)
por: Kawahara, Hideki, et al.
Publicado: (2024)
Determined blind source separation via modeling adjacent frequency band correlations in speech signals
por: Wang, Jianyu, et al.
Publicado: (2025)
por: Wang, Jianyu, et al.
Publicado: (2025)
EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones
por: Hauret, Julien, et al.
Publicado: (2022)
por: Hauret, Julien, et al.
Publicado: (2022)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
por: Tabatabaee, Saba, et al.
Publicado: (2026)
por: Tabatabaee, Saba, et al.
Publicado: (2026)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
por: Honda, Tomoki, et al.
Publicado: (2024)
por: Honda, Tomoki, et al.
Publicado: (2024)
On the relationship between speech and hearing
por: Umesh, Srinivasan, et al.
Publicado: (2024)
por: Umesh, Srinivasan, et al.
Publicado: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
por: Ducorroy, Alexandre, et al.
Publicado: (2025)
por: Ducorroy, Alexandre, et al.
Publicado: (2025)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
por: Deng, Qingkun, et al.
Publicado: (2024)
por: Deng, Qingkun, et al.
Publicado: (2024)
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
por: Shi, Hao, et al.
Publicado: (2025)
por: Shi, Hao, et al.
Publicado: (2025)
Probing mental health information in speech foundation models
por: de Gennes, Marc, et al.
Publicado: (2024)
por: de Gennes, Marc, et al.
Publicado: (2024)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
por: Li, Junjie, et al.
Publicado: (2024)
por: Li, Junjie, et al.
Publicado: (2024)
Distilling a speech and music encoder with task arithmetic
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2025)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2025)
WhisperFlow: speech foundation models in real time
por: Wang, Rongxiang, et al.
Publicado: (2024)
por: Wang, Rongxiang, et al.
Publicado: (2024)
A methodological framework and exemplar protocol for the collection and analysis of repeated speech samples
por: Cummins, Nicholas, et al.
Publicado: (2024)
por: Cummins, Nicholas, et al.
Publicado: (2024)
Omni-directional attention mechanism based on Mamba for speech separation
por: Xue, Ke, et al.
Publicado: (2026)
por: Xue, Ke, et al.
Publicado: (2026)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
por: Ohnaka, Hien, et al.
Publicado: (2024)
por: Ohnaka, Hien, et al.
Publicado: (2024)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
por: Rehman, Abdul, et al.
Publicado: (2025)
por: Rehman, Abdul, et al.
Publicado: (2025)
SPGM: Prioritizing Local Features for enhanced speech separation performance
por: Yip, Jia Qi, et al.
Publicado: (2023)
por: Yip, Jia Qi, et al.
Publicado: (2023)
Inter-channel Conv-TasNet for multichannel speech enhancement
por: Lee, Dongheon, et al.
Publicado: (2021)
por: Lee, Dongheon, et al.
Publicado: (2021)
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
por: Gao, Yuan, et al.
Publicado: (2025)
por: Gao, Yuan, et al.
Publicado: (2025)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
por: Yanir, Efrayim, et al.
Publicado: (2025)
por: Yanir, Efrayim, et al.
Publicado: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
por: Liu, Xueyu, et al.
Publicado: (2024)
por: Liu, Xueyu, et al.
Publicado: (2024)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
por: Li, Xuyuan, et al.
Publicado: (2023)
por: Li, Xuyuan, et al.
Publicado: (2023)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
por: Pascu, Octavian, et al.
Publicado: (2023)
por: Pascu, Octavian, et al.
Publicado: (2023)
A lightweight and robust method for blind wideband-to-fullband extension of speech
por: Büthe, Jan, et al.
Publicado: (2024)
por: Büthe, Jan, et al.
Publicado: (2024)
Monaural speech enhancement on drone via Adapter based transfer learning
por: Chen, Xingyu, et al.
Publicado: (2024)
por: Chen, Xingyu, et al.
Publicado: (2024)
FreeCodec: A disentangled neural speech codec with fewer tokens
por: Zheng, Youqiang, et al.
Publicado: (2024)
por: Zheng, Youqiang, et al.
Publicado: (2024)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
por: Chary, Podakanti Satyajith
Publicado: (2024)
por: Chary, Podakanti Satyajith
Publicado: (2024)
Adversarial speech for voice privacy protection from Personalized Speech generation
por: Chen, Shihao, et al.
Publicado: (2024)
por: Chen, Shihao, et al.
Publicado: (2024)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
por: Das, Sneha, et al.
Publicado: (2020)
por: Das, Sneha, et al.
Publicado: (2020)
A Comparative Study on Positional Encoding for Time-frequency Domain Dual-path Transformer-based Source Separation Models
por: Saijo, Kohei, et al.
Publicado: (2025)
por: Saijo, Kohei, et al.
Publicado: (2025)
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
por: Biberger, Thomas, et al.
Publicado: (2021)
por: Biberger, Thomas, et al.
Publicado: (2021)
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024)
por: Watanabe, Aya, et al.
Publicado: (2024)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
por: Kammoun, Sofiene, et al.
Publicado: (2025)
por: Kammoun, Sofiene, et al.
Publicado: (2025)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
por: Wang, Jianzong, et al.
Publicado: (2023)
por: Wang, Jianzong, et al.
Publicado: (2023)
Ejemplares similares
-
Sound Safeguarding for Acoustic Measurement Using Any Sounds: Tools and Applications
por: Kawahara, Hideki, et al.
Publicado: (2025) -
Subband Splitting: Simple, Efficient and Effective Technique for Solving Block Permutation Problem in Determined Blind Source Separation
por: Matsumoto, Kazuki, et al.
Publicado: (2024) -
SELEBI: Percussion-aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform
por: Akaishi, Natsuki, et al.
Publicado: (2026) -
Local Equivariance Error-Based Metrics for Evaluating Sampling-Frequency-Independent Property of Neural Network
por: Imamura, Kanami, et al.
Publicado: (2025) -
Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides
por: Imamura, Kanami, et al.
Publicado: (2023)