BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Lekai, Gu, Haoyu, Zhao, Jingwei, Wang, Ziyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
von: Yi, Yungang, et al.
Veröffentlicht: (2024)
von: Yi, Yungang, et al.
Veröffentlicht: (2024)
Step-Audio-R1 Technical Report
von: Tian, Fei, et al.
Veröffentlicht: (2025)
von: Tian, Fei, et al.
Veröffentlicht: (2025)
Generation of Musical Timbres using a Text-Guided Diffusion Model
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2026)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2026)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Improving French Synthetic Speech Quality via SSML Prosody Control
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
von: Ismail, Saifelden M.
Veröffentlicht: (2025)
von: Ismail, Saifelden M.
Veröffentlicht: (2025)
Prevailing Research Areas for Music AI in the Era of Foundation Models
von: Wei, Megan, et al.
Veröffentlicht: (2024)
von: Wei, Megan, et al.
Veröffentlicht: (2024)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Taming Audio VAEs via Target-KL Regularization
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
von: Zhu, Di, et al.
Veröffentlicht: (2026)
von: Zhu, Di, et al.
Veröffentlicht: (2026)
VoiceSHIELD-Small: Real-Time Malicious Speech Detection and Transcription
von: Ranjan, Sumit, et al.
Veröffentlicht: (2026)
von: Ranjan, Sumit, et al.
Veröffentlicht: (2026)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
von: Shu, Hongzhi, et al.
Veröffentlicht: (2024)
von: Shu, Hongzhi, et al.
Veröffentlicht: (2024)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
M6(GPT)3: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time Signature
von: Poćwiardowski, Jakub, et al.
Veröffentlicht: (2024)
von: Poćwiardowski, Jakub, et al.
Veröffentlicht: (2024)
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
von: Chopra, Anuradha, et al.
Veröffentlicht: (2025)
von: Chopra, Anuradha, et al.
Veröffentlicht: (2025)
BMdataset: A Musicologically Curated LilyPond Dataset
von: Spanio, Matteo, et al.
Veröffentlicht: (2026)
von: Spanio, Matteo, et al.
Veröffentlicht: (2026)
To Retrieve or Not to Retrieve? Uncertainty Detection for Dynamic Retrieval Augmented Generation
von: Dhole, Kaustubh D.
Veröffentlicht: (2025)
von: Dhole, Kaustubh D.
Veröffentlicht: (2025)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts
von: Opria, Joshua
Veröffentlicht: (2026)
von: Opria, Joshua
Veröffentlicht: (2026)
VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
von: Basha, Maris, et al.
Veröffentlicht: (2025)
von: Basha, Maris, et al.
Veröffentlicht: (2025)
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
von: Xu, Tianyu, et al.
Veröffentlicht: (2026)
von: Xu, Tianyu, et al.
Veröffentlicht: (2026)
Mixer Metaphors: audio interfaces for non-musical applications
von: McNamara, Tace, et al.
Veröffentlicht: (2025)
von: McNamara, Tace, et al.
Veröffentlicht: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Domain Adaptation of the Pyannote Diarization Pipeline for Conversational Indonesian Audio
von: Prasetyo, Muhammad Daffa'i Rafi, et al.
Veröffentlicht: (2026)
von: Prasetyo, Muhammad Daffa'i Rafi, et al.
Veröffentlicht: (2026)
Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Continuous Latent Contexts Enable Efficient Online Learning in Transformers
von: Anand, Emile, et al.
Veröffentlicht: (2026)
von: Anand, Emile, et al.
Veröffentlicht: (2026)
Masked Contrastive Pre-Training Improves Music Audio Key Detection
von: Yonay, Ori, et al.
Veröffentlicht: (2026)
von: Yonay, Ori, et al.
Veröffentlicht: (2026)
MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface Vibrations for Real-time Noise-Robust Speech Input
von: Hiraki, Hirotaka, et al.
Veröffentlicht: (2025)
von: Hiraki, Hirotaka, et al.
Veröffentlicht: (2025)
Refining music sample identification with a self-supervised graph neural network
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
An audio-to-analysis pipeline with certified transcription for information-theoretic profiling of the piano repertoire
von: Jalbert-Desforges, Fred
Veröffentlicht: (2026)
von: Jalbert-Desforges, Fred
Veröffentlicht: (2026)
Ähnliche Einträge
-
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
von: Wang, Zihao, et al.
Veröffentlicht: (2025) -
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
von: Yi, Yungang, et al.
Veröffentlicht: (2024) -
Step-Audio-R1 Technical Report
von: Tian, Fei, et al.
Veröffentlicht: (2025) -
Generation of Musical Timbres using a Text-Guided Diffusion Model
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025) -
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)