Improving Speech Inversion Through Self-Supervised Embeddings and Enhanced Tract Variables
Fuente:
arXiv
Salvato in:
| Autori principali: | Attia, Ahmed Adel, Siriwardena, Yashish M., Espy-Wilson, Carol |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RealClass: A Framework for Classroom Speech Simulation with Public Datasets and Game Engines
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
A multi-modal approach for identifying schizophrenia using cross-modal attention
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
From Weak Labels to Strong Results: Utilizing 5,000 Hours of Noisy Classroom Transcripts with Minimal Accurate Data
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
Reverse Attention for Lightweight Speech Enhancement on Edge Devices
di: Ojha, Shuubham, et al.
Pubblicazione: (2025)
di: Ojha, Shuubham, et al.
Pubblicazione: (2025)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
di: Medin, Lucas Block, et al.
Pubblicazione: (2025)
di: Medin, Lucas Block, et al.
Pubblicazione: (2025)
Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge
di: Xie, Yuankun, et al.
Pubblicazione: (2024)
di: Xie, Yuankun, et al.
Pubblicazione: (2024)
Semi-Supervised Self-Learning Enhanced Music Emotion Recognition
di: Sun, Yifu, et al.
Pubblicazione: (2024)
di: Sun, Yifu, et al.
Pubblicazione: (2024)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
di: Lachenani, Sidahmed, et al.
Pubblicazione: (2025)
di: Lachenani, Sidahmed, et al.
Pubblicazione: (2025)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
di: Ioannides, Georgios, et al.
Pubblicazione: (2026)
di: Ioannides, Georgios, et al.
Pubblicazione: (2026)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
di: Duret, Jarod, et al.
Pubblicazione: (2024)
di: Duret, Jarod, et al.
Pubblicazione: (2024)
Coding Speech through Vocal Tract Kinematics
di: Cho, Cheol Jun, et al.
Pubblicazione: (2024)
di: Cho, Cheol Jun, et al.
Pubblicazione: (2024)
A Multimodal Framework for the Assessment of the Schizophrenia Spectrum
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
di: Zhang, Hanglei, et al.
Pubblicazione: (2025)
di: Zhang, Hanglei, et al.
Pubblicazione: (2025)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2020)
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2020)
Leveraging Mixture of Experts for Improved Speech Deepfake Detection
di: Negroni, Viola, et al.
Pubblicazione: (2024)
di: Negroni, Viola, et al.
Pubblicazione: (2024)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
di: Hwang, Injune, et al.
Pubblicazione: (2024)
di: Hwang, Injune, et al.
Pubblicazione: (2024)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
di: Dixit, Satvik, et al.
Pubblicazione: (2024)
di: Dixit, Satvik, et al.
Pubblicazione: (2024)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
di: Choi, Yerin, et al.
Pubblicazione: (2024)
di: Choi, Yerin, et al.
Pubblicazione: (2024)
MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
di: Songyi, Li, et al.
Pubblicazione: (2026)
di: Songyi, Li, et al.
Pubblicazione: (2026)
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
di: Alex, Tony, et al.
Pubblicazione: (2025)
di: Alex, Tony, et al.
Pubblicazione: (2025)
Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes
di: Kukanov, Ivan, et al.
Pubblicazione: (2024)
di: Kukanov, Ivan, et al.
Pubblicazione: (2024)
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
di: Das, Shoutrik, et al.
Pubblicazione: (2025)
di: Das, Shoutrik, et al.
Pubblicazione: (2025)
DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
di: Lin, Bin, et al.
Pubblicazione: (2026)
di: Lin, Bin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
RealClass: A Framework for Classroom Speech Simulation with Public Datasets and Game Engines
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025) -
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025) -
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025) -
A multi-modal approach for identifying schizophrenia using cross-modal attention
di: Premananth, Gowtham, et al.
Pubblicazione: (2023) -
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)