Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
Fuente:
arXiv
Guardado en:
| Autores principales: | Engert, Natalie, Wagner, Dominik, Riedhammer, Korbinian, Bocklet, Tobias |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
por: Wagner, Dominik, et al.
Publicado: (2025)
por: Wagner, Dominik, et al.
Publicado: (2025)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
por: Bayerl, Sebastian P., et al.
Publicado: (2022)
por: Bayerl, Sebastian P., et al.
Publicado: (2022)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
por: Baumann, Ilja, et al.
Publicado: (2025)
por: Baumann, Ilja, et al.
Publicado: (2025)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
por: Wagner, Dominik, et al.
Publicado: (2024)
por: Wagner, Dominik, et al.
Publicado: (2024)
Shared Multi-modal Embedding Space for Face-Voice Association
por: Simic, Christopher, et al.
Publicado: (2025)
por: Simic, Christopher, et al.
Publicado: (2025)
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
por: Simic, Christopher, et al.
Publicado: (2025)
por: Simic, Christopher, et al.
Publicado: (2025)
Large Language Models for Dysfluency Detection in Stuttered Speech
por: Wagner, Dominik, et al.
Publicado: (2024)
por: Wagner, Dominik, et al.
Publicado: (2024)
Automatic classification of stop realisation with wav2vec2.0
por: Tanner, James, et al.
Publicado: (2025)
por: Tanner, James, et al.
Publicado: (2025)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
por: Li, Feng, et al.
Publicado: (2024)
por: Li, Feng, et al.
Publicado: (2024)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
por: Wang, Zhiyong, et al.
Publicado: (2024)
por: Wang, Zhiyong, et al.
Publicado: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
por: Guo, Yiwei, et al.
Publicado: (2024)
por: Guo, Yiwei, et al.
Publicado: (2024)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
por: Wagner, Dominik, et al.
Publicado: (2023)
por: Wagner, Dominik, et al.
Publicado: (2023)
Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0
por: Kang, Taein, et al.
Publicado: (2024)
por: Kang, Taein, et al.
Publicado: (2024)
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
por: Zhu, Qiushi, et al.
Publicado: (2024)
por: Zhu, Qiushi, et al.
Publicado: (2024)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
por: Huo, Robin, et al.
Publicado: (2025)
por: Huo, Robin, et al.
Publicado: (2025)
On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
por: Grosman, Jonatas, et al.
Publicado: (2025)
por: Grosman, Jonatas, et al.
Publicado: (2025)
Infusing Acoustic Pause Context into Text-Based Dementia Assessment
por: Braun, Franziska, et al.
Publicado: (2024)
por: Braun, Franziska, et al.
Publicado: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
por: Leung, Wing-Zin, et al.
Publicado: (2024)
por: Leung, Wing-Zin, et al.
Publicado: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
por: Attia, Ahmed Adel, et al.
Publicado: (2024)
por: Attia, Ahmed Adel, et al.
Publicado: (2024)
Optimized Speculative Sampling for GPU Hardware Accelerators
por: Wagner, Dominik, et al.
Publicado: (2024)
por: Wagner, Dominik, et al.
Publicado: (2024)
Probing Whisper for Dysarthric Speech in Detection and Assessment
por: Yue, Zhengjun, et al.
Publicado: (2025)
por: Yue, Zhengjun, et al.
Publicado: (2025)
Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis
por: Wang, Haoshen, et al.
Publicado: (2026)
por: Wang, Haoshen, et al.
Publicado: (2026)
Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
por: Freisinger, Steffen, et al.
Publicado: (2026)
por: Freisinger, Steffen, et al.
Publicado: (2026)
Generalizing to Unseen Disaster Events: A Causal View
por: Seeberger, Philipp, et al.
Publicado: (2025)
por: Seeberger, Philipp, et al.
Publicado: (2025)
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
por: de Groot, Dimme, et al.
Publicado: (2025)
por: de Groot, Dimme, et al.
Publicado: (2025)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
por: Shao, Qijie, et al.
Publicado: (2025)
por: Shao, Qijie, et al.
Publicado: (2025)
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
por: Gulzar, Kashaf, et al.
Publicado: (2025)
por: Gulzar, Kashaf, et al.
Publicado: (2025)
Improved Dysarthric Speech to Text Conversion via TTS Personalization
por: Mihajlik, Péter, et al.
Publicado: (2025)
por: Mihajlik, Péter, et al.
Publicado: (2025)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
por: Wang, Huimeng, et al.
Publicado: (2025)
por: Wang, Huimeng, et al.
Publicado: (2025)
Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios
por: Wang, Shiyao, et al.
Publicado: (2025)
por: Wang, Shiyao, et al.
Publicado: (2025)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
por: Chen, Xueyuan, et al.
Publicado: (2025)
por: Chen, Xueyuan, et al.
Publicado: (2025)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
por: Xie, Xurong, et al.
Publicado: (2022)
por: Xie, Xurong, et al.
Publicado: (2022)
wav2pos: Sound Source Localization using Masked Autoencoders
por: Berg, Axel, et al.
Publicado: (2024)
por: Berg, Axel, et al.
Publicado: (2024)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
por: Lee, Jeehyun, et al.
Publicado: (2024)
por: Lee, Jeehyun, et al.
Publicado: (2024)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
por: Wang, Yuejiao, et al.
Publicado: (2024)
por: Wang, Yuejiao, et al.
Publicado: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
por: Park, Seohyun, et al.
Publicado: (2025)
por: Park, Seohyun, et al.
Publicado: (2025)
Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database
por: Xiao, Qing, et al.
Publicado: (2025)
por: Xiao, Qing, et al.
Publicado: (2025)
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
por: Das, Shoutrik, et al.
Publicado: (2025)
por: Das, Shoutrik, et al.
Publicado: (2025)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
por: Wang, Shiyao, et al.
Publicado: (2025)
por: Wang, Shiyao, et al.
Publicado: (2025)
Ejemplares similares
-
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
por: Wagner, Dominik, et al.
Publicado: (2025) -
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
por: Bayerl, Sebastian P., et al.
Publicado: (2022) -
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
por: Baumann, Ilja, et al.
Publicado: (2025) -
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
por: Wagner, Dominik, et al.
Publicado: (2024) -
Shared Multi-modal Embedding Space for Face-Voice Association
por: Simic, Christopher, et al.
Publicado: (2025)