Distilling Spectrograms into Tokens: Fast and Lightweight Bioacoustic Classification for BirdCLEF+ 2025
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Miyaguchi, Anthony, Gustineli, Murilo, Cheung, Adrian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
Motif Mining and Unsupervised Representation Learning for BirdCLEF 2022
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2022)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2022)
Transfer Learning with Semi-Supervised Dataset Annotation for Birdcall Classification
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2023)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2023)
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
FAST: Fast Audio Spectrogram Transformer
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
von: Rauch, Lukas, et al.
Veröffentlicht: (2024)
von: Rauch, Lukas, et al.
Veröffentlicht: (2024)
Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
von: Xue, Ke, et al.
Veröffentlicht: (2026)
von: Xue, Ke, et al.
Veröffentlicht: (2026)
Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
von: Dehaghania, Parinaz Binandeh, et al.
Veröffentlicht: (2026)
von: Dehaghania, Parinaz Binandeh, et al.
Veröffentlicht: (2026)
ASM: Audio Spectrogram Mixer
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation
von: Soltero, Kaspar, et al.
Veröffentlicht: (2025)
von: Soltero, Kaspar, et al.
Veröffentlicht: (2025)
Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
von: Song, Tianyu, et al.
Veröffentlicht: (2025)
von: Song, Tianyu, et al.
Veröffentlicht: (2025)
Large Language Models and Non-Negative Matrix Factorization for Bioacoustic Signal Decomposition
von: Torabi, Yasaman, et al.
Veröffentlicht: (2025)
von: Torabi, Yasaman, et al.
Veröffentlicht: (2025)
Few-Shot Bioacoustic Event Detection with Frame-Level Embedding Learning System
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
BioME: A Resource-Efficient Bioacoustic Foundational Model for IoT Applications
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2026)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2026)
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
A Practical Guide to Spectrogram Analysis for Audio Signal Processing
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task
von: Phan, Dang Thoai
Veröffentlicht: (2024)
von: Phan, Dang Thoai
Veröffentlicht: (2024)
Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event Detection
von: Chen, Yaxiong, et al.
Veröffentlicht: (2024)
von: Chen, Yaxiong, et al.
Veröffentlicht: (2024)
SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
von: Wan, Zixiang, et al.
Veröffentlicht: (2025)
von: Wan, Zixiang, et al.
Veröffentlicht: (2025)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Deep Space Separable Distillation for Lightweight Acoustic Scene Classification
von: Ye, ShuQi, et al.
Veröffentlicht: (2024)
von: Ye, ShuQi, et al.
Veröffentlicht: (2024)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
Automated Bioacoustic Monitoring for South African Bird Species on Unlabeled Data
von: Doell, Michael, et al.
Veröffentlicht: (2024)
von: Doell, Michael, et al.
Veröffentlicht: (2024)
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Audio Classification of Low Feature Spectrograms Utilizing Convolutional Neural Networks
von: Elias, Noel
Veröffentlicht: (2024)
von: Elias, Noel
Veröffentlicht: (2024)
DQLoRA: A Lightweight Domain-Aware Denoising ASR via Adapter-guided Distillation
von: Yang, Yiru
Veröffentlicht: (2025)
von: Yang, Yiru
Veröffentlicht: (2025)
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Perch 2.0: The Bittern Lesson for Bioacoustics
von: van Merriënboer, Bart, et al.
Veröffentlicht: (2025)
von: van Merriënboer, Bart, et al.
Veröffentlicht: (2025)
Towards Deep Active Learning in Avian Bioacoustics
von: Rauch, Lukas, et al.
Veröffentlicht: (2024)
von: Rauch, Lukas, et al.
Veröffentlicht: (2024)
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification
von: Ritu, Jarin, et al.
Veröffentlicht: (2025)
von: Ritu, Jarin, et al.
Veröffentlicht: (2025)
DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
SELEBI: Percussion-aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform
von: Akaishi, Natsuki, et al.
Veröffentlicht: (2026)
von: Akaishi, Natsuki, et al.
Veröffentlicht: (2026)
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
von: Pozorski, Paweł, et al.
Veröffentlicht: (2026)
von: Pozorski, Paweł, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024) -
Motif Mining and Unsupervised Representation Learning for BirdCLEF 2022
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2022) -
Transfer Learning with Semi-Supervised Dataset Annotation for Birdcall Classification
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2023) -
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024) -
FAST: Fast Audio Spectrogram Transformer
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)