APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Husain, Jaavid Aktar, Herremans, Dorien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prevailing Research Areas for Music AI in the Era of Foundation Models
von: Wei, Megan, et al.
Veröffentlicht: (2024)
von: Wei, Megan, et al.
Veröffentlicht: (2024)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
von: Chopra, Anuradha, et al.
Veröffentlicht: (2025)
von: Chopra, Anuradha, et al.
Veröffentlicht: (2025)
Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
von: Cauzinille, Jules, et al.
Veröffentlicht: (2025)
von: Cauzinille, Jules, et al.
Veröffentlicht: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
von: Combes, Paolo, et al.
Veröffentlicht: (2025)
von: Combes, Paolo, et al.
Veröffentlicht: (2025)
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
von: Du, Wenzhang
Veröffentlicht: (2025)
von: Du, Wenzhang
Veröffentlicht: (2025)
FANAL -- Financial Activity News Alerting Language Modeling Framework
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2025)
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2025)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
von: Khushiyant, et al.
Veröffentlicht: (2026)
von: Khushiyant, et al.
Veröffentlicht: (2026)
CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
von: Cao, Songjun, et al.
Veröffentlicht: (2026)
von: Cao, Songjun, et al.
Veröffentlicht: (2026)
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
von: Gondhalekar, Chinmay, et al.
Veröffentlicht: (2025)
von: Gondhalekar, Chinmay, et al.
Veröffentlicht: (2025)
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
Dense Video Understanding with Gated Residual Tokenization
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Automatic Album Sequencing
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
von: Basha, Maris, et al.
Veröffentlicht: (2025)
von: Basha, Maris, et al.
Veröffentlicht: (2025)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
von: Shu, Hongzhi, et al.
Veröffentlicht: (2024)
von: Shu, Hongzhi, et al.
Veröffentlicht: (2024)
BMdataset: A Musicologically Curated LilyPond Dataset
von: Spanio, Matteo, et al.
Veröffentlicht: (2026)
von: Spanio, Matteo, et al.
Veröffentlicht: (2026)
Quantum-Enhanced Analysis and Grading of Vocal Performance
von: Agarwal, Rohan
Veröffentlicht: (2025)
von: Agarwal, Rohan
Veröffentlicht: (2025)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2025)
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2025)
The Concatenator: A Bayesian Approach To Real Time Concatenative Musaicing
von: Tralie, Christopher, et al.
Veröffentlicht: (2024)
von: Tralie, Christopher, et al.
Veröffentlicht: (2024)
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
Improving French Synthetic Speech Quality via SSML Prosody Control
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
von: He, Zhanhong, et al.
Veröffentlicht: (2025)
von: He, Zhanhong, et al.
Veröffentlicht: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
von: Bell-Navas, Andrés, et al.
Veröffentlicht: (2025)
von: Bell-Navas, Andrés, et al.
Veröffentlicht: (2025)
Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse
von: Meng, Wei
Veröffentlicht: (2025)
von: Meng, Wei
Veröffentlicht: (2025)
Reciprocal Latent Fields for Precomputed Sound Propagation
von: Seuté, Hugo, et al.
Veröffentlicht: (2026)
von: Seuté, Hugo, et al.
Veröffentlicht: (2026)
Visualizing the Evolution of Twitter (X.com) Conversations: A Comprehensive Methodology Applied to AI Training Discussions on ChatGPT
von: Jess, Nicole, et al.
Veröffentlicht: (2024)
von: Jess, Nicole, et al.
Veröffentlicht: (2024)
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
von: Sharma, Akul, et al.
Veröffentlicht: (2025)
von: Sharma, Akul, et al.
Veröffentlicht: (2025)
Evaluating Prompting Strategies for Chart Question Answering with Large Language Models
von: Naikar, Ruthuparna, et al.
Veröffentlicht: (2026)
von: Naikar, Ruthuparna, et al.
Veröffentlicht: (2026)
The evolution of inharmonicity and noisiness in contemporary popular music
von: Deruty, Emmanuel, et al.
Veröffentlicht: (2024)
von: Deruty, Emmanuel, et al.
Veröffentlicht: (2024)
CHORUS: An Agentic Framework for Generating Realistic Deliberation Data
von: Koursaris, A., et al.
Veröffentlicht: (2026)
von: Koursaris, A., et al.
Veröffentlicht: (2026)
Graph Connectionist Temporal Classification for Phoneme Recognition
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
Adaptable Symbolic Music Infilling with MIDI-RWKV
von: Zhou-Zheng, Christian, et al.
Veröffentlicht: (2025)
von: Zhou-Zheng, Christian, et al.
Veröffentlicht: (2025)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Prevailing Research Areas for Music AI in the Era of Foundation Models
von: Wei, Megan, et al.
Veröffentlicht: (2024) -
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025) -
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
von: Chopra, Anuradha, et al.
Veröffentlicht: (2025) -
Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
von: Cauzinille, Jules, et al.
Veröffentlicht: (2025) -
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)