Benchmarking Representations for Speech, Music, and Acoustic Events
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | La Quatra, Moreno, Koudounas, Alkis, Vaiani, Lorenzo, Baralis, Elena, Cagliero, Luca, Garza, Paolo, Siniscalchi, Sabato Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
ITALIC: An Italian Intent Classification Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2023)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2023)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Exploring Generative Error Correction for Dysarthric Speech Recognition
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
MVP: Multi-source Voice Pathology detection
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
An Investigation of Incorporating Mamba for Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2024)
von: Chao, Rong, et al.
Veröffentlicht: (2024)
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
von: Hu, Hu, et al.
Veröffentlicht: (2025)
von: Hu, Hu, et al.
Veröffentlicht: (2025)
Bayesian adaptive learning to latent variables via Variational Bayes and Maximum a Posteriori
von: Hu, Hu, et al.
Veröffentlicht: (2024)
von: Hu, Hu, et al.
Veröffentlicht: (2024)
Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from Speech
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
Speech Analysis of Language Varieties in Italy
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
A Contrastive Learning Approach to Mitigate Bias in Speech Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2025)
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2025)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
A Concept-based approach to Voice Disorder Detection
von: Ghia, Davide, et al.
Veröffentlicht: (2025)
von: Ghia, Davide, et al.
Veröffentlicht: (2025)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Pianoroll-Event: A Novel Score Representation for Symbolic Music
von: Qian, Lekai, et al.
Veröffentlicht: (2026)
von: Qian, Lekai, et al.
Veröffentlicht: (2026)
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
von: Gao, Ming, et al.
Veröffentlicht: (2025)
von: Gao, Ming, et al.
Veröffentlicht: (2025)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
Continual Learning for Acoustic Event Classification
von: Xiao, Yang
Veröffentlicht: (2025)
von: Xiao, Yang
Veröffentlicht: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
Autoregressive Speech Enhancement via Acoustic Tokens
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Voice Disorder Analysis: a Transformer-based Approach
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
Pre-training Autoencoder for Acoustic Event Classification via Blinky
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
Abusive Speech Detection in Indic Languages Using Acoustic Features
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025) -
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025) -
ITALIC: An Italian Intent Classification Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2023) -
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025) -
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)