Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dehaghania, Parinaz Binandeh, Penab, Danilo, Aguiar, A. Pedro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
FAST: Fast Audio Spectrogram Transformer
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
Synthesizer Sound Matching Using Audio Spectrogram Transformers
von: Bruford, Fred, et al.
Veröffentlicht: (2024)
von: Bruford, Fred, et al.
Veröffentlicht: (2024)
Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
Abnormal Respiratory Sound Identification Using Audio-Spectrogram Vision Transformer
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2024)
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2024)
ASM: Audio Spectrogram Mixer
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
Audio Classification of Low Feature Spectrograms Utilizing Convolutional Neural Networks
von: Elias, Noel
Veröffentlicht: (2024)
von: Elias, Noel
Veröffentlicht: (2024)
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
A Practical Guide to Spectrogram Analysis for Audio Signal Processing
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
Multi-View Spectrogram Transformer for Respiratory Sound Classification
von: He, Wentao, et al.
Veröffentlicht: (2023)
von: He, Wentao, et al.
Veröffentlicht: (2023)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
von: Işık, Atakan, et al.
Veröffentlicht: (2025)
von: Işık, Atakan, et al.
Veröffentlicht: (2025)
ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Distilling Spectrograms into Tokens: Fast and Lightweight Bioacoustic Classification for BirdCLEF+ 2025
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2025)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2025)
Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech
von: Huang, Kunyang, et al.
Veröffentlicht: (2025)
von: Huang, Kunyang, et al.
Veröffentlicht: (2025)
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
von: Lee, Dongheon, et al.
Veröffentlicht: (2024)
von: Lee, Dongheon, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
von: Atito, Sara, et al.
Veröffentlicht: (2022)
von: Atito, Sara, et al.
Veröffentlicht: (2022)
Investigation of Feature Selection and Pooling Methods for Environmental Sound Classification
von: Dehaghani, Parinaz Binandeh, et al.
Veröffentlicht: (2025)
von: Dehaghani, Parinaz Binandeh, et al.
Veröffentlicht: (2025)
BAST: Binaural Audio Spectrogram Transformer for Binaural Sound Localization
von: Kuang, Sheng, et al.
Veröffentlicht: (2022)
von: Kuang, Sheng, et al.
Veröffentlicht: (2022)
Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
von: Grau-Haro, Jordi, et al.
Veröffentlicht: (2025)
von: Grau-Haro, Jordi, et al.
Veröffentlicht: (2025)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
Music Genre Classification: A Comparative Analysis of CNN and XGBoost Approaches with Mel-frequency cepstral coefficients and Mel Spectrograms
von: Meng, Yigang
Veröffentlicht: (2024)
von: Meng, Yigang
Veröffentlicht: (2024)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
SELEBI: Percussion-aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform
von: Akaishi, Natsuki, et al.
Veröffentlicht: (2026)
von: Akaishi, Natsuki, et al.
Veröffentlicht: (2026)
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
von: Wang, Pingjie, et al.
Veröffentlicht: (2024)
von: Wang, Pingjie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024) -
FAST: Fast Audio Spectrogram Transformer
von: Naman, Anugunj, et al.
Veröffentlicht: (2025) -
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
von: Schmid, Florian, et al.
Veröffentlicht: (2024) -
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
von: Bae, Sangmin, et al.
Veröffentlicht: (2023) -
Synthesizer Sound Matching Using Audio Spectrogram Transformers
von: Bruford, Fred, et al.
Veröffentlicht: (2024)