Gespeichert in:
| Hauptverfasser: | Ramoneda, Pedro, Parada-Cabaleiro, Emilia, Weck, Benno, Serra, Xavier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.01864 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Explainable and Interpretable Musical Difficulty Estimation: A Parameter-efficient Approach
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
von: Weck, Benno, et al.
Veröffentlicht: (2023)
von: Weck, Benno, et al.
Veröffentlicht: (2023)
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
An Open Research Dataset of the 1932 Cairo Congress of Arab Music
von: Bozkurt, Baris
Veröffentlicht: (2025)
von: Bozkurt, Baris
Veröffentlicht: (2025)
Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2025)
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2025)
Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
InaGVAD : a Challenging French TV and Radio Corpus Annotated for Speech Activity Detection and Speaker Gender Segmentation
von: Doukhan, David, et al.
Veröffentlicht: (2024)
von: Doukhan, David, et al.
Veröffentlicht: (2024)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
Sanidha: A Studio Quality Multi-Modal Dataset for Carnatic Music
von: Krishnan, Venkatakrishnan Vaidyanathapuram, et al.
Veröffentlicht: (2025)
von: Krishnan, Venkatakrishnan Vaidyanathapuram, et al.
Veröffentlicht: (2025)
GraphMuse: A Library for Symbolic Music Graph Processing
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
The language of sound search: Examining User Queries in Audio Search Engines
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
The GigaMIDI Dataset with Features for Expressive Music Performance Detection
von: Lee, Keon Ju Maverick, et al.
Veröffentlicht: (2025)
von: Lee, Keon Ju Maverick, et al.
Veröffentlicht: (2025)
Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search
von: McCallum, Matthew C., et al.
Veröffentlicht: (2024)
von: McCallum, Matthew C., et al.
Veröffentlicht: (2024)
KuiSCIMA v2.0: Improved Baselines, Calibration, and Cross-Notation Generalization for Historical Chinese Music Notations in Jiang Kui's Baishidaoren Gequ
von: Repolusk, Tristan, et al.
Veröffentlicht: (2025)
von: Repolusk, Tristan, et al.
Veröffentlicht: (2025)
Music Proofreading with RefinPaint: Where and How to Modify Compositions given Context
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining
von: Tsoi, Tristan, et al.
Veröffentlicht: (2025)
von: Tsoi, Tristan, et al.
Veröffentlicht: (2025)
The Rest is Silence: Leveraging Unseen Species Models for Computational Musicology
von: Moss, Fabian C., et al.
Veröffentlicht: (2025)
von: Moss, Fabian C., et al.
Veröffentlicht: (2025)
AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
von: Shen, Maohao, et al.
Veröffentlicht: (2024)
von: Shen, Maohao, et al.
Veröffentlicht: (2024)
Style-based Composer Identification and Attribution of Symbolic Music Scores: a Systematic Survey
von: Simonetta, Federico
Veröffentlicht: (2025)
von: Simonetta, Federico
Veröffentlicht: (2025)
WER We Stand: Benchmarking Urdu ASR Models
von: Arif, Samee, et al.
Veröffentlicht: (2024)
von: Arif, Samee, et al.
Veröffentlicht: (2024)
Customizing Speech Recognition Model with Large Language Model Feedback
von: Ling, Shaoshi, et al.
Veröffentlicht: (2025)
von: Ling, Shaoshi, et al.
Veröffentlicht: (2025)
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
von: Robinson, David, et al.
Veröffentlicht: (2024)
von: Robinson, David, et al.
Veröffentlicht: (2024)
How Contrastive Decoding Enhances Large Audio Language Models?
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
Closing the Modality Reasoning Gap for Speech Large Language Models
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
Spatial Audio Processing with Large Language Model on Wearable Devices
von: Mishra, Ayushi, et al.
Veröffentlicht: (2025)
von: Mishra, Ayushi, et al.
Veröffentlicht: (2025)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
SALMONN: Towards Generic Hearing Abilities for Large Language Models
von: Tang, Changli, et al.
Veröffentlicht: (2023)
von: Tang, Changli, et al.
Veröffentlicht: (2023)
Direct Simultaneous Translation Activation for Large Audio-Language Models
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
von: Li, Bohan, et al.
Veröffentlicht: (2025)
von: Li, Bohan, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Explainable and Interpretable Musical Difficulty Estimation: A Parameter-efficient Approach
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024) -
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
von: Weck, Benno, et al.
Veröffentlicht: (2023) -
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
von: Uro, Rémi, et al.
Veröffentlicht: (2024) -
An Open Research Dataset of the 1932 Cairo Congress of Arab Music
von: Bozkurt, Baris
Veröffentlicht: (2025) -
Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2025)