Investigation into respiratory sound classification for an imbalanced data set using hybrid LSTM-KAN architectures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | K. V, Nithinkumar, R, Anand |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset
von: Anastasopoulou, Panagiota, et al.
Veröffentlicht: (2024)
von: Anastasopoulou, Panagiota, et al.
Veröffentlicht: (2024)
EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
von: Chu, Yun, et al.
Veröffentlicht: (2025)
von: Chu, Yun, et al.
Veröffentlicht: (2025)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
PC-MCL: Patient-Consistent Multi-Cycle Learning with multi-label bias correction for respiratory sound classification
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2026)
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2026)
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
von: Borodin, Kirill, et al.
Veröffentlicht: (2024)
von: Borodin, Kirill, et al.
Veröffentlicht: (2024)
Knowledge Distillation for Real-Time Classification of Early Media in Voice Communications
von: Altwlkany, Kemal, et al.
Veröffentlicht: (2024)
von: Altwlkany, Kemal, et al.
Veröffentlicht: (2024)
"I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation
von: Sturm, Bob L. T.
Veröffentlicht: (2025)
von: Sturm, Bob L. T.
Veröffentlicht: (2025)
A Novel Bi-LSTM And Transformer Architecture For Generating Tabla Music
von: Mayya, Roopa, et al.
Veröffentlicht: (2024)
von: Mayya, Roopa, et al.
Veröffentlicht: (2024)
Detecting abnormal heart sound using mobile phones and on-device IConNet
von: Vu, Linh, et al.
Veröffentlicht: (2024)
von: Vu, Linh, et al.
Veröffentlicht: (2024)
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
von: Cui, Hanfang, et al.
Veröffentlicht: (2025)
von: Cui, Hanfang, et al.
Veröffentlicht: (2025)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
von: Villatoro-Tello, Esaú, et al.
Veröffentlicht: (2022)
von: Villatoro-Tello, Esaú, et al.
Veröffentlicht: (2022)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
Enhancing Audio Generation Diversity with Visual Information
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr
von: Yang, Xuwen
Veröffentlicht: (2025)
von: Yang, Xuwen
Veröffentlicht: (2025)
MIRFLEX: Music Information Retrieval Feature Library for Extraction
von: Chopra, Anuradha, et al.
Veröffentlicht: (2024)
von: Chopra, Anuradha, et al.
Veröffentlicht: (2024)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
Synthetic training set generation using text-to-audio models for environmental sound classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
EvMic: Event-based Non-contact sound recovery from effective spatial-temporal modeling
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
Déréverbération non-supervisée de la parole par modèle hybride
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
Towards Enhanced Classification of Abnormal Lung sound in Multi-breath: A Light Weight Multi-label and Multi-head Attention Classification Method
von: Chua, Yi-Wei, et al.
Veröffentlicht: (2024)
von: Chua, Yi-Wei, et al.
Veröffentlicht: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2024)
von: Chao, Rong, et al.
Veröffentlicht: (2024)
Audio Deepfake Attribution: An Initial Dataset and Investigation
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
CycleGuardian: A Framework for Automatic RespiratorySound classification Based on Improved Deep clustering and Contrastive Learning
von: Chu, Yun, et al.
Veröffentlicht: (2025)
von: Chu, Yun, et al.
Veröffentlicht: (2025)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
von: Latifi, Seyed Amir, et al.
Veröffentlicht: (2024)
von: Latifi, Seyed Amir, et al.
Veröffentlicht: (2024)
End-to-End User-Defined Keyword Spotting using Shifted Delta Coefficients
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)
von: Li, Haowen, et al.
Veröffentlicht: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
Binaural sound source localization using a hybrid time and frequency domain model
von: Geva, Gil, et al.
Veröffentlicht: (2024)
von: Geva, Gil, et al.
Veröffentlicht: (2024)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
Ähnliche Einträge
-
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025) -
Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset
von: Anastasopoulou, Panagiota, et al.
Veröffentlicht: (2024) -
EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
von: Chu, Yun, et al.
Veröffentlicht: (2025) -
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025) -
PC-MCL: Patient-Consistent Multi-Cycle Learning with multi-label bias correction for respiratory sound classification
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2026)