"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Koudounas, Alkis, La Quatra, Moreno, Pastor, Eliana, Siniscalchi, Sabato Marco, Baralis, Elena |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
voc2vec: A Foundation Model for Non-Verbal Vocalization
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
Hallucination Benchmark for Speech Foundation Models
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
Exploring Generative Error Correction for Dysarthric Speech Recognition
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
Benchmarking Representations for Speech, Music, and Acoustic Events
di: La Quatra, Moreno, et al.
Pubblicazione: (2024)
di: La Quatra, Moreno, et al.
Pubblicazione: (2024)
MVP: Multi-source Voice Pathology detection
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
A Contrastive Learning Approach to Mitigate Bias in Speech Models
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
ITALIC: An Italian Intent Classification Dataset
di: Koudounas, Alkis, et al.
Pubblicazione: (2023)
di: Koudounas, Alkis, et al.
Pubblicazione: (2023)
Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from Speech
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
Speech Analysis of Language Varieties in Italy
di: La Quatra, Moreno, et al.
Pubblicazione: (2024)
di: La Quatra, Moreno, et al.
Pubblicazione: (2024)
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
di: La Quatra, Moreno, et al.
Pubblicazione: (2024)
di: La Quatra, Moreno, et al.
Pubblicazione: (2024)
Voice Disorder Analysis: a Transformer-based Approach
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
di: Li, Haoyang, et al.
Pubblicazione: (2024)
di: Li, Haoyang, et al.
Pubblicazione: (2024)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
di: Koudounas, Alkis, et al.
Pubblicazione: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
di: Yen, Hao, et al.
Pubblicazione: (2025)
di: Yen, Hao, et al.
Pubblicazione: (2025)
Bayesian adaptive learning to latent variables via Variational Bayes and Maximum a Posteriori
di: Hu, Hu, et al.
Pubblicazione: (2024)
di: Hu, Hu, et al.
Pubblicazione: (2024)
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
di: Yen, Hao, et al.
Pubblicazione: (2026)
di: Yen, Hao, et al.
Pubblicazione: (2026)
A Knowledge-Driven Approach to Target Speech Extraction in the Presence of Background Sound Effects for Cinematic Audio Source Separation (CASS)
di: Ho, Chun-wei, et al.
Pubblicazione: (2026)
di: Ho, Chun-wei, et al.
Pubblicazione: (2026)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
di: Hu, Hu, et al.
Pubblicazione: (2025)
di: Hu, Hu, et al.
Pubblicazione: (2025)
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
di: Ho, Chun-Wei, et al.
Pubblicazione: (2025)
di: Ho, Chun-Wei, et al.
Pubblicazione: (2025)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
A Concept-based approach to Voice Disorder Detection
di: Ghia, Davide, et al.
Pubblicazione: (2025)
di: Ghia, Davide, et al.
Pubblicazione: (2025)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
Continual Contrastive Spoken Language Understanding
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2023)
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2023)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2025)
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2025)
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
di: Everson, Kevin, et al.
Pubblicazione: (2024)
di: Everson, Kevin, et al.
Pubblicazione: (2024)
Acoustic and Semantic Modeling of Emotion in Spoken Language
di: Dutta, Soumya
Pubblicazione: (2026)
di: Dutta, Soumya
Pubblicazione: (2026)
A Knowledge-Driven Approach to Music Segmentation, Music Source Separation and Cinematic Audio Source Separation
di: Ho, Chun-wei, et al.
Pubblicazione: (2026)
di: Ho, Chun-wei, et al.
Pubblicazione: (2026)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
Documenti analoghi
-
voc2vec: A Foundation Model for Non-Verbal Vocalization
di: Koudounas, Alkis, et al.
Pubblicazione: (2025) -
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
di: Koudounas, Alkis, et al.
Pubblicazione: (2025) -
Hallucination Benchmark for Speech Foundation Models
di: Koudounas, Alkis, et al.
Pubblicazione: (2025) -
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
di: Koudounas, Alkis, et al.
Pubblicazione: (2025) -
Exploring Generative Error Correction for Dysarthric Speech Recognition
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)