VoxMed: One-Step Respiratory Disease Classifier using Digital Stethoscope Sounds
Fuente:
arXiv
Salvato in:
| Autori principali: | Mundra, Paridhi, Sharma, Manik, Chaudhuri, Yashwardhan, Phukan, Orchid Chetia, Buduru, Arun Balaji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
di: Chaudhuri, Yashwardhan, et al.
Pubblicazione: (2024)
di: Chaudhuri, Yashwardhan, et al.
Pubblicazione: (2024)
The Reasonable Effectiveness of Speaker Embeddings for Violence Detection
di: Jain, Sarthak, et al.
Pubblicazione: (2024)
di: Jain, Sarthak, et al.
Pubblicazione: (2024)
AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
di: Sharma, Sarthak, et al.
Pubblicazione: (2024)
di: Sharma, Sarthak, et al.
Pubblicazione: (2024)
PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
di: Koshal, Devyani, et al.
Pubblicazione: (2024)
di: Koshal, Devyani, et al.
Pubblicazione: (2024)
Towards Neural Audio Codec Source Parsing
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
ComFeAT: Combination of Neural and Spectral Features for Improved Depression Detection
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
di: Jain, Sarthak, et al.
Pubblicazione: (2024)
di: Jain, Sarthak, et al.
Pubblicazione: (2024)
PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
di: Girish, et al.
Pubblicazione: (2025)
di: Girish, et al.
Pubblicazione: (2025)
CoLLAB: A Collaborative Approach for Multilingual Abuse Detection
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Towards Machine Unlearning for Paralinguistic Speech Processing
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2025)
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2025)
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages
di: Girish, et al.
Pubblicazione: (2026)
di: Girish, et al.
Pubblicazione: (2026)
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Towards Multilingual Audio-Visual Question Answering
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions
di: Koo, Heejoon, et al.
Pubblicazione: (2026)
di: Koo, Heejoon, et al.
Pubblicazione: (2026)
Rethinking Cross-Corpus Speech Emotion Recognition Benchmarking: Are Paralinguistic Pre-Trained Representations Sufficient?
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Adaptive Differential Denoising for Respiratory Sounds Classification
di: Dong, Gaoyang, et al.
Pubblicazione: (2025)
di: Dong, Gaoyang, et al.
Pubblicazione: (2025)
Are Multimodal Foundation Models All That Is Needed for Emofake Detection?
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2025)
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2025)
Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
di: Wei, Peidong, et al.
Pubblicazione: (2025)
di: Wei, Peidong, et al.
Pubblicazione: (2025)
NeuRO: An Application for Code-Switched Autism Detection in Children
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2024)
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2024)
HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
di: Kacprzak, Stanisław, et al.
Pubblicazione: (2024)
di: Kacprzak, Stanisław, et al.
Pubblicazione: (2024)
DiffVox: A Differentiable Model for Capturing and Analysing Vocal Effects Distributions
di: Yu, Chin-Yun, et al.
Pubblicazione: (2025)
di: Yu, Chin-Yun, et al.
Pubblicazione: (2025)
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection
di: Yu, Hogeon
Pubblicazione: (2025)
di: Yu, Hogeon
Pubblicazione: (2025)
A Lightweight Feature Fusion Architecture For Resource-Constrained Crowd Counting
di: Chaudhuri, Yashwardhan, et al.
Pubblicazione: (2024)
di: Chaudhuri, Yashwardhan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
di: Chaudhuri, Yashwardhan, et al.
Pubblicazione: (2024) -
The Reasonable Effectiveness of Speaker Embeddings for Violence Detection
di: Jain, Sarthak, et al.
Pubblicazione: (2024) -
AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
di: Sharma, Sarthak, et al.
Pubblicazione: (2024) -
PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
di: Koshal, Devyani, et al.
Pubblicazione: (2024) -
Towards Neural Audio Codec Source Parsing
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)