Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Niizumi, Daisuke, Takeuchi, Daiki, Yasuda, Masahiro, Nguyen, Binh Thien, Ohishi, Yasunori, Harada, Noboru |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Pre-training an Effective Respiratory Audio Foundation Model
by: Niizumi, Daisuke, et al.
Published: (2025)
by: Niizumi, Daisuke, et al.
Published: (2025)
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
by: Niizumi, Daisuke, et al.
Published: (2025)
by: Niizumi, Daisuke, et al.
Published: (2025)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
by: Niizumi, Daisuke, et al.
Published: (2026)
by: Niizumi, Daisuke, et al.
Published: (2026)
Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
by: Nguyen, Binh Thien, et al.
Published: (2025)
by: Nguyen, Binh Thien, et al.
Published: (2025)
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
by: Takeuchi, Daiki, et al.
Published: (2025)
by: Takeuchi, Daiki, et al.
Published: (2025)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
by: Tsubaki, Shunsuke, et al.
Published: (2024)
by: Tsubaki, Shunsuke, et al.
Published: (2024)
Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources
by: Nguyen, Binh Thien, et al.
Published: (2026)
by: Nguyen, Binh Thien, et al.
Published: (2026)
Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
by: Yasuda, Masahiro, et al.
Published: (2025)
by: Yasuda, Masahiro, et al.
Published: (2025)
Description and Discussion on DCASE 2026 Challenge Task 2: Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
by: Nishida, Tomoya, et al.
Published: (2026)
by: Nishida, Tomoya, et al.
Published: (2026)
Can Sound Replace Vision in LLaVA With Token Substitution?
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on self-motioning human
by: Yasuda, Masahiro, et al.
Published: (2024)
by: Yasuda, Masahiro, et al.
Published: (2024)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
by: Kawamura, Takao, et al.
Published: (2026)
by: Kawamura, Takao, et al.
Published: (2026)
Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN
by: Zhang, Shiqi, et al.
Published: (2024)
by: Zhang, Shiqi, et al.
Published: (2024)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
by: Guichoux, Téo, et al.
Published: (2025)
by: Guichoux, Téo, et al.
Published: (2025)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
Description and Discussion on DCASE 2025 Challenge Task 2: First-shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
by: Nishida, Tomoya, et al.
Published: (2025)
by: Nishida, Tomoya, et al.
Published: (2025)
SARA: Stress Test Reasoning in Audio Deepfake Detection
by: Nguyen, Binh, et al.
Published: (2026)
by: Nguyen, Binh, et al.
Published: (2026)
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
by: Salehi, Pegah, et al.
Published: (2024)
by: Salehi, Pegah, et al.
Published: (2024)
FakeSound: Deepfake General Audio Detection
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
by: Nguyen, Binh Thien, et al.
Published: (2026)
by: Nguyen, Binh Thien, et al.
Published: (2026)
Description and Discussion on DCASE 2024 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
by: Nishida, Tomoya, et al.
Published: (2024)
by: Nishida, Tomoya, et al.
Published: (2024)
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
by: Wang, Pingjie, et al.
Published: (2024)
by: Wang, Pingjie, et al.
Published: (2024)
Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and Localization
by: Wu, Junyan, et al.
Published: (2024)
by: Wu, Junyan, et al.
Published: (2024)
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
by: Roy, Abhinaba, et al.
Published: (2025)
by: Roy, Abhinaba, et al.
Published: (2025)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
BowelRCNN: Region-based Convolutional Neural Network System for Bowel Sound Auscultation
by: Matynia, Igor, et al.
Published: (2025)
by: Matynia, Igor, et al.
Published: (2025)
Sound Safeguarding for Acoustic Measurement Using Any Sounds: Tools and Applications
by: Kawahara, Hideki, et al.
Published: (2025)
by: Kawahara, Hideki, et al.
Published: (2025)
Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event Analysis
by: Yasuda, Masahiro, et al.
Published: (2024)
by: Yasuda, Masahiro, et al.
Published: (2024)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Adaptive Differential Denoising for Respiratory Sounds Classification
by: Dong, Gaoyang, et al.
Published: (2025)
by: Dong, Gaoyang, et al.
Published: (2025)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
by: Zheng, Zihao, et al.
Published: (2025)
by: Zheng, Zihao, et al.
Published: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
A Speech Enhancement Method Using Fast Fourier Transform and Convolutional Autoencoder
by: Kow, Pu-Yun, et al.
Published: (2025)
by: Kow, Pu-Yun, et al.
Published: (2025)
Abnormal Respiratory Sound Identification Using Audio-Spectrogram Vision Transformer
by: Ariyanti, Whenty, et al.
Published: (2024)
by: Ariyanti, Whenty, et al.
Published: (2024)
Acousto-optic reconstruction of exterior sound field based on concentric circle sampling with circular harmonic expansion
by: Nguyen, Phuc Duc, et al.
Published: (2023)
by: Nguyen, Phuc Duc, et al.
Published: (2023)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
by: Sarkar, Eklavya, et al.
Published: (2024)
by: Sarkar, Eklavya, et al.
Published: (2024)
Similar Items
-
Towards Pre-training an Effective Respiratory Audio Foundation Model
by: Niizumi, Daisuke, et al.
Published: (2025) -
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
by: Niizumi, Daisuke, et al.
Published: (2024) -
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
by: Niizumi, Daisuke, et al.
Published: (2024) -
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
by: Niizumi, Daisuke, et al.
Published: (2024) -
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
by: Niizumi, Daisuke, et al.
Published: (2025)