SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zihan, Ding, Shuangrui, Zhang, Zhixiong, Dong, Xiaoyi, Zhang, Pan, Zang, Yuhang, Cao, Yuhang, Lin, Dahua, Wang, Jiaqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
von: Li, Jiajia, et al.
Veröffentlicht: (2026)
von: Li, Jiajia, et al.
Veröffentlicht: (2026)
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
von: Qian, Rui, et al.
Veröffentlicht: (2025)
von: Qian, Rui, et al.
Veröffentlicht: (2025)
SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation
von: Li, Sifei, et al.
Veröffentlicht: (2026)
von: Li, Sifei, et al.
Veröffentlicht: (2026)
Advancing Complex Video Object Segmentation via Progressive Concept Construction
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
SongCreator: Lyrics-based Universal Song Generation
von: Lei, Shun, et al.
Veröffentlicht: (2024)
von: Lei, Shun, et al.
Veröffentlicht: (2024)
SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control
von: Jiang, Changhao, et al.
Veröffentlicht: (2026)
von: Jiang, Changhao, et al.
Veröffentlicht: (2026)
SegTune: Structured and Fine-Grained Control for Song Generation
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
Versatile Framework for Song Generation with Prompt-based Control
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
LeVo: High-Quality Song Generation with Multi-Preference Alignment
von: Lei, Shun, et al.
Veröffentlicht: (2025)
von: Lei, Shun, et al.
Veröffentlicht: (2025)
Analyzing Pitch Content in Traditional Ghanaian Seperewa Songs
von: Walls, Kelvin L, et al.
Veröffentlicht: (2024)
von: Walls, Kelvin L, et al.
Veröffentlicht: (2024)
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
von: Bu, Jiazi, et al.
Veröffentlicht: (2024)
von: Bu, Jiazi, et al.
Veröffentlicht: (2024)
AI-Generated Song Detection via Lyrics Transcripts
von: Frohmann, Markus, et al.
Veröffentlicht: (2025)
von: Frohmann, Markus, et al.
Veröffentlicht: (2025)
The Florence Price Art Song Dataset and Piano Accompaniment Generator
von: He, Tao-Tao, et al.
Veröffentlicht: (2025)
von: He, Tao-Tao, et al.
Veröffentlicht: (2025)
An End-to-End Approach for Chord-Conditioned Song Generation
von: Gao, Shuochen, et al.
Veröffentlicht: (2024)
von: Gao, Shuochen, et al.
Veröffentlicht: (2024)
Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart Parsing
von: Li, Jinsong, et al.
Veröffentlicht: (2026)
von: Li, Jinsong, et al.
Veröffentlicht: (2026)
Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models
von: Li, Jinsong, et al.
Veröffentlicht: (2025)
von: Li, Jinsong, et al.
Veröffentlicht: (2025)
PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation
von: Ou, Longshen, et al.
Veröffentlicht: (2025)
von: Ou, Longshen, et al.
Veröffentlicht: (2025)
Come Together: Analyzing Popular Songs Through Statistical Embeddings
von: Mallory, Matthew Esmaili, et al.
Veröffentlicht: (2026)
von: Mallory, Matthew Esmaili, et al.
Veröffentlicht: (2026)
SongTrans: An unified song transcription and alignment method for lyrics and notes
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2026)
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2026)
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation
von: Zhang, Huaicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Huaicheng, et al.
Veröffentlicht: (2025)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
von: Chen, Huakang, et al.
Veröffentlicht: (2025)
von: Chen, Huakang, et al.
Veröffentlicht: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
von: Cao, Yuhang, et al.
Veröffentlicht: (2024)
von: Cao, Yuhang, et al.
Veröffentlicht: (2024)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
von: Liu, Renhang, et al.
Veröffentlicht: (2025)
von: Liu, Renhang, et al.
Veröffentlicht: (2025)
SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-Training
von: Yu, Jiaxing, et al.
Veröffentlicht: (2024)
von: Yu, Jiaxing, et al.
Veröffentlicht: (2024)
SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers
von: Chen, Guangke, et al.
Veröffentlicht: (2024)
von: Chen, Guangke, et al.
Veröffentlicht: (2024)
TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
von: He, Qi, et al.
Veröffentlicht: (2025)
von: He, Qi, et al.
Veröffentlicht: (2025)
Segment-Factorized Full-Song Generation on Symbolic Piano Music
von: Chen, Ping-Yi, et al.
Veröffentlicht: (2025)
von: Chen, Ping-Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024) -
2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025) -
SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
von: Li, Jiajia, et al.
Veröffentlicht: (2026) -
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
von: Qian, Rui, et al.
Veröffentlicht: (2025) -
SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation
von: Li, Sifei, et al.
Veröffentlicht: (2026)