A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Oli Danyi, Tang, Hao, Feldman, Naomi, Goldwater, Sharon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Effective Context in Neural Speech Models
von: Meng, Yen, et al.
Veröffentlicht: (2025)
von: Meng, Yen, et al.
Veröffentlicht: (2025)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
von: Yanir, Efrayim, et al.
Veröffentlicht: (2025)
von: Yanir, Efrayim, et al.
Veröffentlicht: (2025)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
In-context learning capabilities of Large Language Models to detect suicide risk among adolescents from speech transcripts
von: Roquefort, Filomene, et al.
Veröffentlicht: (2025)
von: Roquefort, Filomene, et al.
Veröffentlicht: (2025)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
FreeCodec: A disentangled neural speech codec with fewer tokens
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
von: Ma, Lu
Veröffentlicht: (2025)
von: Ma, Lu
Veröffentlicht: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
Monaural speech enhancement on drone via Adapter based transfer learning
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
Probing mental health information in speech foundation models
von: de Gennes, Marc, et al.
Veröffentlicht: (2024)
von: de Gennes, Marc, et al.
Veröffentlicht: (2024)
WhisperFlow: speech foundation models in real time
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
Semantic enrichment towards efficient speech representations
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
von: Biberger, Thomas, et al.
Veröffentlicht: (2021)
von: Biberger, Thomas, et al.
Veröffentlicht: (2021)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
von: Samui, Suman
Veröffentlicht: (2022)
von: Samui, Suman
Veröffentlicht: (2022)
Multichannel blind speech source separation with a disjoint constraint source model
von: Wang, Jianyu, et al.
Veröffentlicht: (2024)
von: Wang, Jianyu, et al.
Veröffentlicht: (2024)
Disentangling peripheral hearing loss from central and cognitive effects on speech intelligibility in older adults
von: Irino, Toshio, et al.
Veröffentlicht: (2025)
von: Irino, Toshio, et al.
Veröffentlicht: (2025)
Self-supervised learning of speech representations with Dutch archival data
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
On the relationship between speech and hearing
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
An automatic analysis of ultrasound vocalisations for the prediction of interaction context in captive Egyptian fruit bats
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
von: Wang, Yi, et al.
Veröffentlicht: (2025) -
Effective Context in Neural Speech Models
von: Meng, Yen, et al.
Veröffentlicht: (2025) -
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
von: Yanir, Efrayim, et al.
Veröffentlicht: (2025) -
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024) -
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)