Persian Musical Instruments Classification Using Polyphonic Data Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Esfangereh, Diba Hadi, Sameti, Mohammad Hossein, Moridani, Sepehr Harfi, Javidpour, Leili, Baghshah, Mahdieh Soleymani |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2026)
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2026)
Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark
von: Kalahroodi, Mohammad Javad Ranjbar, et al.
Veröffentlicht: (2026)
von: Kalahroodi, Mohammad Javad Ranjbar, et al.
Veröffentlicht: (2026)
A Neural Score Follower for Computer Accompaniment of Polyphonic Musical Instruments
von: Pillay, Ashwin
Veröffentlicht: (2025)
von: Pillay, Ashwin
Veröffentlicht: (2025)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
von: Rashidi, Sina, et al.
Veröffentlicht: (2025)
von: Rashidi, Sina, et al.
Veröffentlicht: (2025)
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
von: Paqaleh, Mohammad Mahdi Samiei, et al.
Veröffentlicht: (2025)
von: Paqaleh, Mohammad Mahdi Samiei, et al.
Veröffentlicht: (2025)
Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
von: Rahmani, Zahra, et al.
Veröffentlicht: (2025)
von: Rahmani, Zahra, et al.
Veröffentlicht: (2025)
Do Music Preferences Reflect Cultural Values? A Cross-National Analysis Using Music Embedding and World Values Survey
von: Kim, Yongjae, et al.
Veröffentlicht: (2025)
von: Kim, Yongjae, et al.
Veröffentlicht: (2025)
ELAB: Extensive LLM Alignment Benchmark in Persian Language
von: Pourbahman, Zahra, et al.
Veröffentlicht: (2025)
von: Pourbahman, Zahra, et al.
Veröffentlicht: (2025)
Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey
von: Shahhosseini, Fatemeh, et al.
Veröffentlicht: (2025)
von: Shahhosseini, Fatemeh, et al.
Veröffentlicht: (2025)
Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR
von: Lee, Minsik, et al.
Veröffentlicht: (2026)
von: Lee, Minsik, et al.
Veröffentlicht: (2026)
No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
BASS: Benchmarking Audio LMs for Musical Structure and Semantic Reasoning
von: Jang, Min, et al.
Veröffentlicht: (2026)
von: Jang, Min, et al.
Veröffentlicht: (2026)
Story2MIDI: Emotionally Aligned Music Generation from Text
von: Shokri, Mohammad, et al.
Veröffentlicht: (2025)
von: Shokri, Mohammad, et al.
Veröffentlicht: (2025)
Rubato: Transcribing Piano Music with Timestamps
von: Tamer, Nazif Can, et al.
Veröffentlicht: (2026)
von: Tamer, Nazif Can, et al.
Veröffentlicht: (2026)
From Image to Music Language: A Two-Stage Structure Decoding Approach for Complex Polyphonic OMR
von: Xu, Nan, et al.
Veröffentlicht: (2026)
von: Xu, Nan, et al.
Veröffentlicht: (2026)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
von: Weck, Benno, et al.
Veröffentlicht: (2026)
von: Weck, Benno, et al.
Veröffentlicht: (2026)
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
von: Li, Jiajia, et al.
Veröffentlicht: (2026)
von: Li, Jiajia, et al.
Veröffentlicht: (2026)
Transcribing Rhythmic Patterns of the Guitar Track in Polyphonic Music
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2025)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2025)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
von: Sedghiyeh, Nima, et al.
Veröffentlicht: (2025)
von: Sedghiyeh, Nima, et al.
Veröffentlicht: (2025)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
von: Wei, Haojie, et al.
Veröffentlicht: (2023)
von: Wei, Haojie, et al.
Veröffentlicht: (2023)
Adaptive Test-Time Scaling for Zero-Shot Respiratory Audio Classification
von: Wang, Tsai-Ning, et al.
Veröffentlicht: (2026)
von: Wang, Tsai-Ning, et al.
Veröffentlicht: (2026)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition
von: Ramezani, Erfan, et al.
Veröffentlicht: (2026)
von: Ramezani, Erfan, et al.
Veröffentlicht: (2026)
MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core
von: Liao, Callie C., et al.
Veröffentlicht: (2025)
von: Liao, Callie C., et al.
Veröffentlicht: (2025)
Assessing Factual Music Comprehension in Large Audio Language Models
von: Lin, Daniel Chenyu, et al.
Veröffentlicht: (2025)
von: Lin, Daniel Chenyu, et al.
Veröffentlicht: (2025)
Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean
von: Joo, Hyunjung, et al.
Veröffentlicht: (2026)
von: Joo, Hyunjung, et al.
Veröffentlicht: (2026)
CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2026) -
Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025) -
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025) -
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024) -
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)