Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre Classification
Fuente:
arXiv
Saved in:
| Main Author: | Sridhar, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Music Genre Classification using Large Language Models
by: Meguenani, Mohamed El Amine, et al.
Published: (2024)
by: Meguenani, Mohamed El Amine, et al.
Published: (2024)
The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
by: Namgyal, Tashi, et al.
Published: (2024)
by: Namgyal, Tashi, et al.
Published: (2024)
Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings
by: Gohari, Mahyar, et al.
Published: (2024)
by: Gohari, Mahyar, et al.
Published: (2024)
Multi-View Spectrogram Transformer for Respiratory Sound Classification
by: He, Wentao, et al.
Published: (2023)
by: He, Wentao, et al.
Published: (2023)
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
by: Atito, Sara, et al.
Published: (2022)
by: Atito, Sara, et al.
Published: (2022)
ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
by: Gupta, Pranav, et al.
Published: (2024)
by: Gupta, Pranav, et al.
Published: (2024)
Music Genre Classification: Training an AI model
by: Mogonediwa, Keoikantse
Published: (2024)
by: Mogonediwa, Keoikantse
Published: (2024)
A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions
by: Ma, Junwen, et al.
Published: (2026)
by: Ma, Junwen, et al.
Published: (2026)
GCDance: Genre-Controlled Music-Driven 3D Full Body Dance Generation
by: Liu, Xinran, et al.
Published: (2025)
by: Liu, Xinran, et al.
Published: (2025)
Dynamic Cross Attention for Audio-Visual Person Verification
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Acoustic Scene Classification: A Competition Review
by: Gharib, Shayan, et al.
Published: (2018)
by: Gharib, Shayan, et al.
Published: (2018)
Audio Processing using Pattern Recognition for Music Genre Classification
by: Chatterjee, Sivangi, et al.
Published: (2024)
by: Chatterjee, Sivangi, et al.
Published: (2024)
Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
by: Guragain, Anmol
Published: (2026)
by: Guragain, Anmol
Published: (2026)
PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation
by: Im, Jaekwon, et al.
Published: (2026)
by: Im, Jaekwon, et al.
Published: (2026)
Synthesizing Audio from Silent Video using Sequence to Sequence Modeling
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2024)
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2024)
V2Meow: Meowing to the Visual Beat via Video-to-Music Generation
by: Su, Kun, et al.
Published: (2023)
by: Su, Kun, et al.
Published: (2023)
Convolutional Neural Network Achieves Human-level Accuracy in Music Genre Classification
by: Dong, Mingwen
Published: (2018)
by: Dong, Mingwen
Published: (2018)
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
by: Tang, Mingni, et al.
Published: (2025)
by: Tang, Mingni, et al.
Published: (2025)
Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
by: Zhu, Yongxin, et al.
Published: (2024)
by: Zhu, Yongxin, et al.
Published: (2024)
Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation
by: Zeng, Runhao, et al.
Published: (2025)
by: Zeng, Runhao, et al.
Published: (2025)
MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation
by: Takahashi, Akira, et al.
Published: (2026)
by: Takahashi, Akira, et al.
Published: (2026)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
by: Takahashi, Akira, et al.
Published: (2025)
by: Takahashi, Akira, et al.
Published: (2025)
When Vision Models Meet Parameter Efficient Look-Aside Adapters Without Large-Scale Audio Pretraining
by: Yeo, Juan, et al.
Published: (2024)
by: Yeo, Juan, et al.
Published: (2024)
Tune It Up: Music Genre Transfer and Prediction
by: Samet, Fidan, et al.
Published: (2025)
by: Samet, Fidan, et al.
Published: (2025)
FLUX that Plays Music
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
Music Genre Classification: Ensemble Learning with Subcomponents-level Attention
by: Liu, Yichen, et al.
Published: (2024)
by: Liu, Yichen, et al.
Published: (2024)
Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
by: Ríos-Vila, Antonio, et al.
Published: (2024)
by: Ríos-Vila, Antonio, et al.
Published: (2024)
Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
by: Sun, Changchang, et al.
Published: (2025)
by: Sun, Changchang, et al.
Published: (2025)
UniMuMo: Unified Text, Music and Motion Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
Unboxing Engagement in YouTube Influencer Videos: An Attention-Based Approach
by: Rajaram, Prashant, et al.
Published: (2020)
by: Rajaram, Prashant, et al.
Published: (2020)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
Assessing the Robustness of Spectral Clustering for Deep Speaker Diarization
by: Raghav, Nikhil, et al.
Published: (2024)
by: Raghav, Nikhil, et al.
Published: (2024)
Exploring Green AI for Audio Deepfake Detection
by: Saha, Subhajit, et al.
Published: (2024)
by: Saha, Subhajit, et al.
Published: (2024)
Joint Multimodal Transformer for Emotion Recognition in the Wild
by: Waligora, Paul, et al.
Published: (2024)
by: Waligora, Paul, et al.
Published: (2024)
Dynamic Modality and View Selection for Multimodal Emotion Recognition with Missing Modalities
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
by: Malard, Hugo, et al.
Published: (2024)
by: Malard, Hugo, et al.
Published: (2024)
Developing an AI-based Integrated System for Bee Health Evaluation
by: Liang, Andrew
Published: (2024)
by: Liang, Andrew
Published: (2024)
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
by: Cheng, Luyao, et al.
Published: (2024)
by: Cheng, Luyao, et al.
Published: (2024)
Similar Items
-
Music Genre Classification using Large Language Models
by: Meguenani, Mohamed El Amine, et al.
Published: (2024) -
The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
by: Namgyal, Tashi, et al.
Published: (2024) -
Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings
by: Gohari, Mahyar, et al.
Published: (2024) -
Multi-View Spectrogram Transformer for Respiratory Sound Classification
by: He, Wentao, et al.
Published: (2023) -
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)