DRAGON: Distributional Rewards Optimize Diffusion Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Yatong, Casebeer, Jonah, Sojoudi, Somayeh, Bryan, Nicholas J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
von: Ma, Yinghao, et al.
Veröffentlicht: (2026)
von: Ma, Yinghao, et al.
Veröffentlicht: (2026)
FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
von: Fan, Pingyi, et al.
Veröffentlicht: (2025)
von: Fan, Pingyi, et al.
Veröffentlicht: (2025)
Efficient Fine-Grained Guidance for Diffusion Model Based Symbolic Music Generation
von: Zhu, Tingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Tingyu, et al.
Veröffentlicht: (2024)
JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models
von: Li, Peike, et al.
Veröffentlicht: (2023)
von: Li, Peike, et al.
Veröffentlicht: (2023)
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models
von: Mehta, Atharva, et al.
Veröffentlicht: (2025)
von: Mehta, Atharva, et al.
Veröffentlicht: (2025)
Stemphonic: All-at-once Flexible Multi-stem Music Generation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2026)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2026)
Presto! Distilling Steps and Layers for Accelerating Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
von: Sun, Yirong, et al.
Veröffentlicht: (2025)
von: Sun, Yirong, et al.
Veröffentlicht: (2025)
D3PIA: A Discrete Denoising Diffusion Model for Piano Accompaniment Generation From Lead sheet
von: Choi, Eunjin, et al.
Veröffentlicht: (2026)
von: Choi, Eunjin, et al.
Veröffentlicht: (2026)
MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
von: Boudaghi, Ali, et al.
Veröffentlicht: (2025)
von: Boudaghi, Ali, et al.
Veröffentlicht: (2025)
Generative AI for Music and Audio
von: Dong, Hao-Wen
Veröffentlicht: (2024)
von: Dong, Hao-Wen
Veröffentlicht: (2024)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
von: Shao, Keren, et al.
Veröffentlicht: (2025)
von: Shao, Keren, et al.
Veröffentlicht: (2025)
Fast Text-to-Audio Generation with Adversarial Post-Training
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
Segment-Factorized Full-Song Generation on Symbolic Piano Music
von: Chen, Ping-Yi, et al.
Veröffentlicht: (2025)
von: Chen, Ping-Yi, et al.
Veröffentlicht: (2025)
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation
von: Nagarajan, Ashwin, et al.
Veröffentlicht: (2025)
von: Nagarajan, Ashwin, et al.
Veröffentlicht: (2025)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
Automatic Time Signature Determination for New Scores Using Lyrics for Latent Rhythmic Structure
von: Liao, Callie C., et al.
Veröffentlicht: (2023)
von: Liao, Callie C., et al.
Veröffentlicht: (2023)
Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Video-based Music Generation
von: Sulun, Serkan
Veröffentlicht: (2026)
von: Sulun, Serkan
Veröffentlicht: (2026)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026)
von: Li, Haitian, et al.
Veröffentlicht: (2026)
S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
von: Hong, Jiaying, et al.
Veröffentlicht: (2025)
von: Hong, Jiaying, et al.
Veröffentlicht: (2025)
Automatic Music Transcription using Convolutional Neural Networks and Constant-Q transform
von: Telila, Yohannis, et al.
Veröffentlicht: (2025)
von: Telila, Yohannis, et al.
Veröffentlicht: (2025)
From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2025)
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2025)
On the de-duplication of the Lakh MIDI dataset
von: Choi, Eunjin, et al.
Veröffentlicht: (2025)
von: Choi, Eunjin, et al.
Veröffentlicht: (2025)
Audio Transformers
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
Sequence-to-Sequence Multi-Modal Speech In-Painting
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
von: Natesan, Sanjay, et al.
Veröffentlicht: (2024)
von: Natesan, Sanjay, et al.
Veröffentlicht: (2024)
Content Adaptive Front End For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
von: Bai, Yatong, et al.
Veröffentlicht: (2023) -
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026) -
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
von: Novack, Zachary, et al.
Veröffentlicht: (2026) -
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025) -
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
von: Ma, Yinghao, et al.
Veröffentlicht: (2026)