Probing mental health information in speech foundation models
Fuente:
arXiv
Salvato in:
| Autori principali: | de Gennes, Marc, Lesage, Adrien, Denais, Martin, Cao, Xuan-Nga, Chang, Simon, Van Remoortere, Pierre, Dakhlia, Cyrille, Riad, Rachid |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
di: Ando, Angelika, et al.
Pubblicazione: (2025)
di: Ando, Angelika, et al.
Pubblicazione: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Quantized Approximate Signal Processing (QASP): Towards Homomorphic Encryption for audio
di: Nguyen, Tu Duyen, et al.
Pubblicazione: (2025)
di: Nguyen, Tu Duyen, et al.
Pubblicazione: (2025)
In-context learning capabilities of Large Language Models to detect suicide risk among adolescents from speech transcripts
di: Roquefort, Filomene, et al.
Pubblicazione: (2025)
di: Roquefort, Filomene, et al.
Pubblicazione: (2025)
WhisperFlow: speech foundation models in real time
di: Wang, Rongxiang, et al.
Pubblicazione: (2024)
di: Wang, Rongxiang, et al.
Pubblicazione: (2024)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
A lightweight and robust method for blind wideband-to-fullband extension of speech
di: Büthe, Jan, et al.
Pubblicazione: (2024)
di: Büthe, Jan, et al.
Pubblicazione: (2024)
Omni-directional attention mechanism based on Mamba for speech separation
di: Xue, Ke, et al.
Pubblicazione: (2026)
di: Xue, Ke, et al.
Pubblicazione: (2026)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
Deep low-latency joint speech transmission and enhancement over a gaussian channel
di: Bokaei, Mohammad, et al.
Pubblicazione: (2024)
di: Bokaei, Mohammad, et al.
Pubblicazione: (2024)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
On the relationship between speech and hearing
di: Umesh, Srinivasan, et al.
Pubblicazione: (2024)
di: Umesh, Srinivasan, et al.
Pubblicazione: (2024)
Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speech
di: Dineley, Judith, et al.
Pubblicazione: (2023)
di: Dineley, Judith, et al.
Pubblicazione: (2023)
The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement
di: Leglaive, Simon, et al.
Pubblicazione: (2023)
di: Leglaive, Simon, et al.
Pubblicazione: (2023)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
di: Li, Junjie, et al.
Pubblicazione: (2024)
di: Li, Junjie, et al.
Pubblicazione: (2024)
Distilling a speech and music encoder with task arithmetic
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
di: Ohnaka, Hien, et al.
Pubblicazione: (2024)
di: Ohnaka, Hien, et al.
Pubblicazione: (2024)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
SPGM: Prioritizing Local Features for enhanced speech separation performance
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
Inter-channel Conv-TasNet for multichannel speech enhancement
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
Probing Self-supervised Learning Models with Target Speech Extraction
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
di: Liu, Xueyu, et al.
Pubblicazione: (2024)
di: Liu, Xueyu, et al.
Pubblicazione: (2024)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
di: Li, Xuyuan, et al.
Pubblicazione: (2023)
di: Li, Xuyuan, et al.
Pubblicazione: (2023)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
Monaural speech enhancement on drone via Adapter based transfer learning
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
FreeCodec: A disentangled neural speech codec with fewer tokens
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
di: Chary, Podakanti Satyajith
Pubblicazione: (2024)
di: Chary, Podakanti Satyajith
Pubblicazione: (2024)
Adversarial speech for voice privacy protection from Personalized Speech generation
di: Chen, Shihao, et al.
Pubblicazione: (2024)
di: Chen, Shihao, et al.
Pubblicazione: (2024)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
di: Das, Sneha, et al.
Pubblicazione: (2020)
di: Das, Sneha, et al.
Pubblicazione: (2020)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
di: Lyth, Dan, et al.
Pubblicazione: (2024)
di: Lyth, Dan, et al.
Pubblicazione: (2024)
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
di: Biberger, Thomas, et al.
Pubblicazione: (2021)
di: Biberger, Thomas, et al.
Pubblicazione: (2021)
Building speech corpus with diverse voice characteristics for its prompt-based representation
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
Language model integration based on memory control for sequence to sequence speech recognition
di: Cho, Jaejin, et al.
Pubblicazione: (2018)
di: Cho, Jaejin, et al.
Pubblicazione: (2018)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
di: Ma, Te, et al.
Pubblicazione: (2025)
di: Ma, Te, et al.
Pubblicazione: (2025)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
di: Samui, Suman
Pubblicazione: (2022)
di: Samui, Suman
Pubblicazione: (2022)
Documenti analoghi
-
SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
di: Ando, Angelika, et al.
Pubblicazione: (2025) -
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025) -
Quantized Approximate Signal Processing (QASP): Towards Homomorphic Encryption for audio
di: Nguyen, Tu Duyen, et al.
Pubblicazione: (2025) -
In-context learning capabilities of Large Language Models to detect suicide risk among adolescents from speech transcripts
di: Roquefort, Filomene, et al.
Pubblicazione: (2025) -
WhisperFlow: speech foundation models in real time
di: Wang, Rongxiang, et al.
Pubblicazione: (2024)