SongTrans: An unified song transcription and alignment method for lyrics and notes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Siwei, He, Jinzheng, Yuan, Ruibin, Wei, Haojie, Wei, Xipin, Lin, Chenghua, Xu, Jin, Lin, Junyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
von: Tan, Wei, et al.
Veröffentlicht: (2025)
von: Tan, Wei, et al.
Veröffentlicht: (2025)
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2026)
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2026)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
The Florence Price Art Song Dataset and Piano Accompaniment Generator
von: He, Tao-Tao, et al.
Veröffentlicht: (2025)
von: He, Tao-Tao, et al.
Veröffentlicht: (2025)
Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
von: Hou, Siyuan, et al.
Veröffentlicht: (2024)
von: Hou, Siyuan, et al.
Veröffentlicht: (2024)
Analyzing and Mitigating Inconsistency in Discrete Audio Tokens for Neural Codec Language Models
von: Liu, Wenrui, et al.
Veröffentlicht: (2024)
von: Liu, Wenrui, et al.
Veröffentlicht: (2024)
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
von: Wei, Haojie, et al.
Veröffentlicht: (2023)
von: Wei, Haojie, et al.
Veröffentlicht: (2023)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
von: Wei, Haojie, et al.
Veröffentlicht: (2024)
von: Wei, Haojie, et al.
Veröffentlicht: (2024)
LeVo: High-Quality Song Generation with Multi-Preference Alignment
von: Lei, Shun, et al.
Veröffentlicht: (2025)
von: Lei, Shun, et al.
Veröffentlicht: (2025)
Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
von: Wei, Peidong, et al.
Veröffentlicht: (2025)
von: Wei, Peidong, et al.
Veröffentlicht: (2025)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
von: Chen, Huakang, et al.
Veröffentlicht: (2025)
von: Chen, Huakang, et al.
Veröffentlicht: (2025)
FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
SongCreator: Lyrics-based Universal Song Generation
von: Lei, Shun, et al.
Veröffentlicht: (2024)
von: Lei, Shun, et al.
Veröffentlicht: (2024)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
Automatic Live Music Song Identification Using Multi-level Deep Sequence Similarity Learning
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation
von: Wei, Haojie, et al.
Veröffentlicht: (2025)
von: Wei, Haojie, et al.
Veröffentlicht: (2025)
DashengTokenizer: One layer is enough for unified audio understanding and generation
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Joint sentiment analysis of lyrics and audio in music
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
DGSNA: Dynamic Generative Scene-based Noise Addition method
von: Chen, Zihao, et al.
Veröffentlicht: (2024)
von: Chen, Zihao, et al.
Veröffentlicht: (2024)
Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
von: Riley, Xavier, et al.
Veröffentlicht: (2024)
von: Riley, Xavier, et al.
Veröffentlicht: (2024)
Qwen2-Audio Technical Report
von: Chu, Yunfei, et al.
Veröffentlicht: (2024)
von: Chu, Yunfei, et al.
Veröffentlicht: (2024)
TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
von: He, Qi, et al.
Veröffentlicht: (2025)
von: He, Qi, et al.
Veröffentlicht: (2025)
Peransformer: Improving Low-informed Expressive Performance Rendering with Score-aware Discriminator
von: He, Xian, et al.
Veröffentlicht: (2025)
von: He, Xian, et al.
Veröffentlicht: (2025)
Multi-label Cross-lingual automatic music genre classification from lyrics with Sentence BERT
von: Tavares, Tiago Fernandes, et al.
Veröffentlicht: (2025)
von: Tavares, Tiago Fernandes, et al.
Veröffentlicht: (2025)
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
ICGAN: An implicit conditioning method for interpretable feature control of neural audio synthesis
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
von: Tan, Wei, et al.
Veröffentlicht: (2025) -
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
von: Yang, Chenyu, et al.
Veröffentlicht: (2024) -
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2026) -
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023) -
The Florence Price Art Song Dataset and Piano Accompaniment Generator
von: He, Tao-Tao, et al.
Veröffentlicht: (2025)