Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pešán, Jan, Kesiraju, Santosh, Burget, Lukáš, Černocký, Jan ''Honza'' |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
Robustness of Speech Separation Models for Similar-pitch Speakers
von: Lay, Bunlong, et al.
Veröffentlicht: (2024)
von: Lay, Bunlong, et al.
Veröffentlicht: (2024)
Unsupervised Speech Enhancement using Data-defined Priors
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
Quartered Chirp Spectral Envelope for Whispered vs Normal Speech Classification
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech
von: Mohammad, Mir Sayeed, et al.
Veröffentlicht: (2024)
von: Mohammad, Mir Sayeed, et al.
Veröffentlicht: (2024)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
von: Serbest, Sanberk, et al.
Veröffentlicht: (2025)
von: Serbest, Sanberk, et al.
Veröffentlicht: (2025)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
BUT System for the MLC-SLM Challenge
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
Speech Watermarking with Discrete Intermediate Representations
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
von: Gupta, Kishan, et al.
Veröffentlicht: (2022)
von: Gupta, Kishan, et al.
Veröffentlicht: (2022)
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
Lightweight DNN for Full-Band Speech Denoising on Mobile Devices: Exploiting Long and Short Temporal Patterns
von: Drossos, Konstantinos, et al.
Veröffentlicht: (2025)
von: Drossos, Konstantinos, et al.
Veröffentlicht: (2025)
State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
von: Barahona, Sara, et al.
Veröffentlicht: (2024)
von: Barahona, Sara, et al.
Veröffentlicht: (2024)
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
AI-Assisted Music Production: A User Study on Text-to-Music Models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
von: Sun, Esther, et al.
Veröffentlicht: (2026)
von: Sun, Esther, et al.
Veröffentlicht: (2026)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Semantic Communications for Speech Recognition
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
von: Marmor, Yanir, et al.
Veröffentlicht: (2026)
von: Marmor, Yanir, et al.
Veröffentlicht: (2026)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
von: Storey, Edward, et al.
Veröffentlicht: (2025)
von: Storey, Edward, et al.
Veröffentlicht: (2025)
Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality
von: Marinoni, Christian, et al.
Veröffentlicht: (2024)
von: Marinoni, Christian, et al.
Veröffentlicht: (2024)
Comparison of Tiny Machine Learning Techniques for Embedded Acoustic Emission Analysis
von: Muthumala, Uditha, et al.
Veröffentlicht: (2024)
von: Muthumala, Uditha, et al.
Veröffentlicht: (2024)
Is Audio Spoof Detection Robust to Laundering Attacks?
von: Ali, Hashim, et al.
Veröffentlicht: (2024)
von: Ali, Hashim, et al.
Veröffentlicht: (2024)
Ultra Low Complexity Deep Learning Based Noise Suppression
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2023)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
von: Polok, Alexander, et al.
Veröffentlicht: (2024) -
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2025) -
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2026) -
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025) -
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)