Investigation into respiratory sound classification for an imbalanced data set using hybrid LSTM-KAN architectures
Fuente:
arXiv
Salvato in:
| Autori principali: | K. V, Nithinkumar, R, Anand |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
di: Emon, Jakaria Islam, et al.
Pubblicazione: (2025)
di: Emon, Jakaria Islam, et al.
Pubblicazione: (2025)
Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset
di: Anastasopoulou, Panagiota, et al.
Pubblicazione: (2024)
di: Anastasopoulou, Panagiota, et al.
Pubblicazione: (2024)
EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
di: Chu, Yun, et al.
Pubblicazione: (2025)
di: Chu, Yun, et al.
Pubblicazione: (2025)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
di: Li, Yuqi, et al.
Pubblicazione: (2025)
di: Li, Yuqi, et al.
Pubblicazione: (2025)
PC-MCL: Patient-Consistent Multi-Cycle Learning with multi-label bias correction for respiratory sound classification
di: Jeong, Seung Gyu, et al.
Pubblicazione: (2026)
di: Jeong, Seung Gyu, et al.
Pubblicazione: (2026)
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
di: Olvera, Michel, et al.
Pubblicazione: (2024)
di: Olvera, Michel, et al.
Pubblicazione: (2024)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
di: Borodin, Kirill, et al.
Pubblicazione: (2024)
di: Borodin, Kirill, et al.
Pubblicazione: (2024)
Knowledge Distillation for Real-Time Classification of Early Media in Voice Communications
di: Altwlkany, Kemal, et al.
Pubblicazione: (2024)
di: Altwlkany, Kemal, et al.
Pubblicazione: (2024)
"I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation
di: Sturm, Bob L. T.
Pubblicazione: (2025)
di: Sturm, Bob L. T.
Pubblicazione: (2025)
A Novel Bi-LSTM And Transformer Architecture For Generating Tabla Music
di: Mayya, Roopa, et al.
Pubblicazione: (2024)
di: Mayya, Roopa, et al.
Pubblicazione: (2024)
Detecting abnormal heart sound using mobile phones and on-device IConNet
di: Vu, Linh, et al.
Pubblicazione: (2024)
di: Vu, Linh, et al.
Pubblicazione: (2024)
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
di: Cui, Hanfang, et al.
Pubblicazione: (2025)
di: Cui, Hanfang, et al.
Pubblicazione: (2025)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
di: Mehta, Shivam, et al.
Pubblicazione: (2025)
di: Mehta, Shivam, et al.
Pubblicazione: (2025)
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
di: Villatoro-Tello, Esaú, et al.
Pubblicazione: (2022)
di: Villatoro-Tello, Esaú, et al.
Pubblicazione: (2022)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
di: Yang, Xiaoran, et al.
Pubblicazione: (2024)
di: Yang, Xiaoran, et al.
Pubblicazione: (2024)
Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model
di: Oluwademilade, Adelekun, et al.
Pubblicazione: (2026)
di: Oluwademilade, Adelekun, et al.
Pubblicazione: (2026)
Enhancing Audio Generation Diversity with Visual Information
di: Xie, Zeyu, et al.
Pubblicazione: (2024)
di: Xie, Zeyu, et al.
Pubblicazione: (2024)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr
di: Yang, Xuwen
Pubblicazione: (2025)
di: Yang, Xuwen
Pubblicazione: (2025)
MIRFLEX: Music Information Retrieval Feature Library for Extraction
di: Chopra, Anuradha, et al.
Pubblicazione: (2024)
di: Chopra, Anuradha, et al.
Pubblicazione: (2024)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
di: Li, Kai, et al.
Pubblicazione: (2024)
di: Li, Kai, et al.
Pubblicazione: (2024)
Synthetic training set generation using text-to-audio models for environmental sound classification
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
EvMic: Event-based Non-contact sound recovery from effective spatial-temporal modeling
di: Yin, Hao, et al.
Pubblicazione: (2025)
di: Yin, Hao, et al.
Pubblicazione: (2025)
Déréverbération non-supervisée de la parole par modèle hybride
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
Towards Enhanced Classification of Abnormal Lung sound in Multi-breath: A Light Weight Multi-label and Multi-head Attention Classification Method
di: Chua, Yi-Wei, et al.
Pubblicazione: (2024)
di: Chua, Yi-Wei, et al.
Pubblicazione: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
Audio Deepfake Attribution: An Initial Dataset and Investigation
di: Yan, Xinrui, et al.
Pubblicazione: (2022)
di: Yan, Xinrui, et al.
Pubblicazione: (2022)
CycleGuardian: A Framework for Automatic RespiratorySound classification Based on Improved Deep clustering and Contrastive Learning
di: Chu, Yun, et al.
Pubblicazione: (2025)
di: Chu, Yun, et al.
Pubblicazione: (2025)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
di: Barański, Mateusz, et al.
Pubblicazione: (2025)
di: Barański, Mateusz, et al.
Pubblicazione: (2025)
Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
di: Latifi, Seyed Amir, et al.
Pubblicazione: (2024)
di: Latifi, Seyed Amir, et al.
Pubblicazione: (2024)
End-to-End User-Defined Keyword Spotting using Shifted Delta Coefficients
di: V, Kesavaraj, et al.
Pubblicazione: (2024)
di: V, Kesavaraj, et al.
Pubblicazione: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
di: Cheng, Zhuangfei, et al.
Pubblicazione: (2025)
di: Cheng, Zhuangfei, et al.
Pubblicazione: (2025)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
di: Li, Haowen, et al.
Pubblicazione: (2025)
di: Li, Haowen, et al.
Pubblicazione: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
di: Halpern, Bence Mark, et al.
Pubblicazione: (2024)
di: Halpern, Bence Mark, et al.
Pubblicazione: (2024)
Binaural sound source localization using a hybrid time and frequency domain model
di: Geva, Gil, et al.
Pubblicazione: (2024)
di: Geva, Gil, et al.
Pubblicazione: (2024)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
di: Donepudi, Dharma Teja
Pubblicazione: (2025)
di: Donepudi, Dharma Teja
Pubblicazione: (2025)
Documenti analoghi
-
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
di: Emon, Jakaria Islam, et al.
Pubblicazione: (2025) -
Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset
di: Anastasopoulou, Panagiota, et al.
Pubblicazione: (2024) -
EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
di: Chu, Yun, et al.
Pubblicazione: (2025) -
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
di: Li, Yuqi, et al.
Pubblicazione: (2025) -
PC-MCL: Patient-Consistent Multi-Cycle Learning with multi-label bias correction for respiratory sound classification
di: Jeong, Seung Gyu, et al.
Pubblicazione: (2026)