MARBLE: Multi-Aspect Reward Balance for Diffusion RL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Canyu, Chen, Hao, Tong, Yunze, Qiao, Yu, Li, Jiacheng, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
Noise Projection: Closing the Prompt-Agnostic Gap Behind Text-to-Image Misalignment in Diffusion Models
von: Tong, Yunze, et al.
Veröffentlicht: (2025)
von: Tong, Yunze, et al.
Veröffentlicht: (2025)
On the Trajectory Regularity of ODE-based Diffusion Sampling
von: Chen, Defang, et al.
Veröffentlicht: (2024)
von: Chen, Defang, et al.
Veröffentlicht: (2024)
A Geometric Perspective on Diffusion Models
von: Chen, Defang, et al.
Veröffentlicht: (2023)
von: Chen, Defang, et al.
Veröffentlicht: (2023)
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
von: Wang, Yiyang, et al.
Veröffentlicht: (2026)
von: Wang, Yiyang, et al.
Veröffentlicht: (2026)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
von: Li, Liyang, et al.
Veröffentlicht: (2026)
von: Li, Liyang, et al.
Veröffentlicht: (2026)
Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning
von: Xie, Desai, et al.
Veröffentlicht: (2023)
von: Xie, Desai, et al.
Veröffentlicht: (2023)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
von: Wu, Junyi, et al.
Veröffentlicht: (2026)
von: Wu, Junyi, et al.
Veröffentlicht: (2026)
LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning
von: Zhang, Mingyang, et al.
Veröffentlicht: (2023)
von: Zhang, Mingyang, et al.
Veröffentlicht: (2023)
RL-Selector: Reinforcement Learning-Guided Data Selection via Redundancy Assessment
von: Yang, Suorong, et al.
Veröffentlicht: (2025)
von: Yang, Suorong, et al.
Veröffentlicht: (2025)
Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
von: Zhang, Ziyi, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2024)
FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
von: Ding, Ganggui, et al.
Veröffentlicht: (2024)
von: Ding, Ganggui, et al.
Veröffentlicht: (2024)
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
von: Huang, Tao, et al.
Veröffentlicht: (2023)
von: Huang, Tao, et al.
Veröffentlicht: (2023)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
Exploring 3D Dataset Pruning
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
Generalization of Diffusion Models Arises with a Balanced Representation Space
von: Zhang, Zekai, et al.
Veröffentlicht: (2025)
von: Zhang, Zekai, et al.
Veröffentlicht: (2025)
Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
BiGain: Unified Token Compression for Joint Generation and Classification
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
Directly Fine-Tuning Diffusion Models on Differentiable Rewards
von: Clark, Kevin, et al.
Veröffentlicht: (2023)
von: Clark, Kevin, et al.
Veröffentlicht: (2023)
FOD-Diff: 3D Multi-Channel Patch Diffusion Model for Fiber Orientation Distribution
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
von: Yang, Kai, et al.
Veröffentlicht: (2023)
von: Yang, Kai, et al.
Veröffentlicht: (2023)
Enhancing Medical Image Segmentation with Deep Learning and Diffusion Models
von: Liu, Houze, et al.
Veröffentlicht: (2024)
von: Liu, Houze, et al.
Veröffentlicht: (2024)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
PromptRL: Prompt Matters in RL for Flow-Based Image Generation
von: Wang, Fu-Yun, et al.
Veröffentlicht: (2026)
von: Wang, Fu-Yun, et al.
Veröffentlicht: (2026)
Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
Pixel-wise RL on Diffusion Models: Reinforcement Learning from Rich Feedback
von: Kordzanganeh, Mo, et al.
Veröffentlicht: (2024)
von: Kordzanganeh, Mo, et al.
Veröffentlicht: (2024)
VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
von: Dai, Fengyuan, et al.
Veröffentlicht: (2025)
von: Dai, Fengyuan, et al.
Veröffentlicht: (2025)
FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
von: Fu, Yuwei, et al.
Veröffentlicht: (2024)
von: Fu, Yuwei, et al.
Veröffentlicht: (2024)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
Adaptive Transformer Attention and Multi-Scale Fusion for Spine 3D Segmentation
von: Xiang, Yanlin, et al.
Veröffentlicht: (2025)
von: Xiang, Yanlin, et al.
Veröffentlicht: (2025)
Mitigating Accuracy-Robustness Trade-off via Balanced Multi-Teacher Adversarial Distillation
von: Zhao, Shiji, et al.
Veröffentlicht: (2023)
von: Zhao, Shiji, et al.
Veröffentlicht: (2023)
DiffPattern-Flex: Efficient Layout Pattern Generation via Discrete Diffusion
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization
von: Wang, Xiaohan, et al.
Veröffentlicht: (2025)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2025)
Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
von: Wang, Austin, et al.
Veröffentlicht: (2026)
von: Wang, Austin, et al.
Veröffentlicht: (2026)
ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control
von: Golan, Shelly, et al.
Veröffentlicht: (2026)
von: Golan, Shelly, et al.
Veröffentlicht: (2026)
Balancing Multimodal Domain Generalization via Gradient Modulation and Projection
von: Li, Hongzhao, et al.
Veröffentlicht: (2026)
von: Li, Hongzhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
von: Zhao, Canyu, et al.
Veröffentlicht: (2025) -
Noise Projection: Closing the Prompt-Agnostic Gap Behind Text-to-Image Misalignment in Diffusion Models
von: Tong, Yunze, et al.
Veröffentlicht: (2025) -
On the Trajectory Regularity of ODE-based Diffusion Sampling
von: Chen, Defang, et al.
Veröffentlicht: (2024) -
A Geometric Perspective on Diffusion Models
von: Chen, Defang, et al.
Veröffentlicht: (2023) -
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
von: Wang, Yiyang, et al.
Veröffentlicht: (2026)