ILD-VIT: A Unified Vision Transformer Architecture for Detection of Interstitial Lung Disease from Respiratory Sounds
Fuente:
arXiv
Saved in:
| Main Authors: | Hota, Soubhagya Ranjan, Roy, Arka, Satija, Udit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers
by: Mang, Loredana Daria, et al.
Published: (2024)
by: Mang, Loredana Daria, et al.
Published: (2024)
XAI-Driven Spectral Analysis of Cough Sounds for Respiratory Disease Characterization
by: Amado-Caballero, Patricia, et al.
Published: (2025)
by: Amado-Caballero, Patricia, et al.
Published: (2025)
Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
by: Jiang, Ya, et al.
Published: (2024)
by: Jiang, Ya, et al.
Published: (2024)
Detection of manatee vocalisations using the Audio Spectrogram Transformer
by: Schiappacasse, Stefano, et al.
Published: (2024)
by: Schiappacasse, Stefano, et al.
Published: (2024)
STNet: Prediction of Underwater Sound Speed Profiles with An Advanced Semi-Transformer Neural Network
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
by: Dumpis, Martynas, et al.
Published: (2026)
by: Dumpis, Martynas, et al.
Published: (2026)
A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction
by: Damiano, Stefano, et al.
Published: (2024)
by: Damiano, Stefano, et al.
Published: (2024)
MASSLOC: A Massive Sound Source Localization System based on Direction-of-Arrival Estimation
by: Fischer, Georg K. J., et al.
Published: (2025)
by: Fischer, Georg K. J., et al.
Published: (2025)
Sound field estimation with moving microphones using kernel ridge regression
by: Brunnström, Jesper, et al.
Published: (2025)
by: Brunnström, Jesper, et al.
Published: (2025)
Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
by: Bhattacharjee, Sankha Subhra, et al.
Published: (2024)
by: Bhattacharjee, Sankha Subhra, et al.
Published: (2024)
SUNAC: Source-aware Unified Neural Audio Codec
by: Aihara, Ryo, et al.
Published: (2025)
by: Aihara, Ryo, et al.
Published: (2025)
FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Hybrid SMI Realization via Matrix Completion and Riemannian Manifold Optimization on Narrowband Sub-Array Based Architectures
by: Cousik, Tarun Suman, et al.
Published: (2026)
by: Cousik, Tarun Suman, et al.
Published: (2026)
Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
String Sound Synthesizer on GPU-accelerated Finite Difference Scheme
by: Lee, Jin Woo, et al.
Published: (2023)
by: Lee, Jin Woo, et al.
Published: (2023)
Reduce Computational Complexity for Continuous Wavelet Transform in Acoustic Recognition Using Hop Size
by: Phan, Dang Thoai
Published: (2024)
by: Phan, Dang Thoai
Published: (2024)
On the Invariance of Cross-Correlation Peak Positions Under Monotonic Signal Transformations, with Application to Fast Time Difference Estimation
by: Ueno, Natsuki, et al.
Published: (2025)
by: Ueno, Natsuki, et al.
Published: (2025)
Soundscape Captioning using Sound Affective Quality Network and Large Language Model
by: Hou, Yuanbo, et al.
Published: (2024)
by: Hou, Yuanbo, et al.
Published: (2024)
Online Similarity-and-Independence-Aware Beamformer for Low-latency Target Sound Extraction
by: Hiroe, Atsuo
Published: (2023)
by: Hiroe, Atsuo
Published: (2023)
Large Language Model-based Nonnegative Matrix Factorization For Cardiorespiratory Sound Separation
by: Torabi, Yasaman, et al.
Published: (2025)
by: Torabi, Yasaman, et al.
Published: (2025)
Benchmarking Audio Deepfake Detection Robustness in Real-world Communication Scenarios
by: Shi, Haohan, et al.
Published: (2025)
by: Shi, Haohan, et al.
Published: (2025)
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
by: Neri, Michael, et al.
Published: (2025)
by: Neri, Michael, et al.
Published: (2025)
Multi-channel Replay Speech Detection using an Adaptive Learnable Beamformer
by: Neri, Michael, et al.
Published: (2025)
by: Neri, Michael, et al.
Published: (2025)
Chirp Group Delay based Onset Detection in Instruments with Fast Attack
by: Joysingh, S. Johanan, et al.
Published: (2024)
by: Joysingh, S. Johanan, et al.
Published: (2024)
Advanced Signal Analysis in Detecting Replay Attacks for Automatic Speaker Verification Systems
by: Kuang, Lee Shih
Published: (2024)
by: Kuang, Lee Shih
Published: (2024)
SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
by: Yao, Shengshi, et al.
Published: (2025)
by: Yao, Shengshi, et al.
Published: (2025)
Robust Detection of Underwater Target Against Non-Uniform Noise With Optical Fiber DAS Array
by: Cang, Siyuan, et al.
Published: (2025)
by: Cang, Siyuan, et al.
Published: (2025)
Decomposing the Influence of Physical Acoustic Modeling on Neural Personal Sound Zone Rendering: An Ablation Study
by: Jiang, Hao, et al.
Published: (2026)
by: Jiang, Hao, et al.
Published: (2026)
FasTUSS: Faster Task-Aware Unified Source Separation
by: Paissan, Francesco, et al.
Published: (2025)
by: Paissan, Francesco, et al.
Published: (2025)
A Multimodal Data Fusion Attention-Empowered Generative Adversarial Network for Real Time 3D Underwater Sound Speed Field Construction
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
by: Wang, Ziqian, et al.
Published: (2025)
by: Wang, Ziqian, et al.
Published: (2025)
A XAI-based Framework for Frequency Subband Characterization of Cough Spectrograms in Chronic Respiratory Disease
by: Amado-Caballero, Patricia, et al.
Published: (2025)
by: Amado-Caballero, Patricia, et al.
Published: (2025)
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech
by: Mohammad, Mir Sayeed, et al.
Published: (2024)
by: Mohammad, Mir Sayeed, et al.
Published: (2024)
WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection
by: Xuan, Xi, et al.
Published: (2026)
by: Xuan, Xi, et al.
Published: (2026)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
by: Du, Hui-Peng, et al.
Published: (2024)
by: Du, Hui-Peng, et al.
Published: (2024)
Auditory Attention Decoding from Ear-EEG Signals: A Dataset with Dynamic Attention Switching and Rigorous Cross-Validation
by: Zhang, Yuanming, et al.
Published: (2025)
by: Zhang, Yuanming, et al.
Published: (2025)
Beyond Identity: A Generalizable Approach for Deepfake Audio Detection
by: Ahmadiadli, Yasaman, et al.
Published: (2025)
by: Ahmadiadli, Yasaman, et al.
Published: (2025)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
by: Kwon, Younghoo, et al.
Published: (2024)
by: Kwon, Younghoo, et al.
Published: (2024)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
by: Aihara, Ryo, et al.
Published: (2025)
by: Aihara, Ryo, et al.
Published: (2025)
Similar Items
-
Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers
by: Mang, Loredana Daria, et al.
Published: (2024) -
XAI-Driven Spectral Analysis of Cough Sounds for Respiratory Disease Characterization
by: Amado-Caballero, Patricia, et al.
Published: (2025) -
Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
by: Jiang, Ya, et al.
Published: (2024) -
Detection of manatee vocalisations using the Audio Spectrogram Transformer
by: Schiappacasse, Stefano, et al.
Published: (2024) -
STNet: Prediction of Underwater Sound Speed Profiles with An Advanced Semi-Transformer Neural Network
by: Huang, Wei, et al.
Published: (2025)