Diversity-Rewarded CFG Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Cideron, Geoffrey, Agostinelli, Andrea, Ferret, Johan, Girgin, Sertan, Elie, Romuald, Bachem, Olivier, Perrin, Sarah, Ramé, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WARM: On the Benefits of Weight Averaged Reward Models
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
WARP: On the Benefits of Weight Averaged Rewarded Policies
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Learning in Mean Field Games: A Survey
by: Laurière, Mathieu, et al.
Published: (2022)
by: Laurière, Mathieu, et al.
Published: (2022)
MusicRL: Aligning Music Generation to Human Preferences
by: Cideron, Geoffrey, et al.
Published: (2024)
by: Cideron, Geoffrey, et al.
Published: (2024)
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Fair Active Learning: Solving the Labeling Problem in Insurance
by: Elie, Romuald, et al.
Published: (2021)
by: Elie, Romuald, et al.
Published: (2021)
Learning Rock Pushability on Rough Planetary Terrain
by: Girgin, Tuba, et al.
Published: (2025)
by: Girgin, Tuba, et al.
Published: (2025)
Contrastive CFG: Improving CFG in Diffusion Models by Contrasting Positive and Negative Concepts
by: Chang, Jinho, et al.
Published: (2024)
by: Chang, Jinho, et al.
Published: (2024)
Fair regression under localized demographic parity constraints
by: Charpentier, Arthur, et al.
Published: (2026)
by: Charpentier, Arthur, et al.
Published: (2026)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
by: Agarwal, Rishabh, et al.
Published: (2023)
by: Agarwal, Rishabh, et al.
Published: (2023)
The DeepXube Software Package for Solving Pathfinding Problems with Learned Heuristic Functions and Search
by: Agostinelli, Forest
Published: (2026)
by: Agostinelli, Forest
Published: (2026)
Clustering in Deep Stochastic Transformers
by: Fedorov, Lev, et al.
Published: (2026)
by: Fedorov, Lev, et al.
Published: (2026)
A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
by: Pignatelli, Eduardo, et al.
Published: (2023)
by: Pignatelli, Eduardo, et al.
Published: (2023)
CFG-OEC: Classifier Free Guidance with Orthogonal Error Correction
by: Yang, Nakgyu, et al.
Published: (2025)
by: Yang, Nakgyu, et al.
Published: (2025)
Guidance in the Frequency Domain Enables High-Fidelity Sampling at Low CFG Scales
by: Sadat, Seyedmorteza, et al.
Published: (2025)
by: Sadat, Seyedmorteza, et al.
Published: (2025)
Control Variate Score Matching for Diffusion Models
by: Kahouli, Khaled, et al.
Published: (2025)
by: Kahouli, Khaled, et al.
Published: (2025)
Optimal Stopping in Latent Diffusion Models
by: Wu, Yu-Han, et al.
Published: (2025)
by: Wu, Yu-Han, et al.
Published: (2025)
Nash Learning from Human Feedback
by: Munos, Rémi, et al.
Published: (2023)
by: Munos, Rémi, et al.
Published: (2023)
Dimension-free error estimate for diffusion model and optimal scheduling
by: de Bortoli, Valentin, et al.
Published: (2025)
by: de Bortoli, Valentin, et al.
Published: (2025)
CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
by: Wang, Hanyang, et al.
Published: (2026)
by: Wang, Hanyang, et al.
Published: (2026)
EP-CFG: Energy-Preserving Classifier-Free Guidance
by: Zhang, Kai, et al.
Published: (2024)
by: Zhang, Kai, et al.
Published: (2024)
MIND: Monge Inception Distance for Generative Models Evaluation
by: Berthet, Quentin, et al.
Published: (2026)
by: Berthet, Quentin, et al.
Published: (2026)
Online Knowledge Distillation with Reward Guidance
by: Jia, Chen
Published: (2025)
by: Jia, Chen
Published: (2025)
CFG++: Manifold-constrained Classifier Free Guidance for Diffusion Models
by: Chung, Hyungjin, et al.
Published: (2024)
by: Chung, Hyungjin, et al.
Published: (2024)
Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings
by: Powell, Keenan, et al.
Published: (2026)
by: Powell, Keenan, et al.
Published: (2026)
Hellinger loss function for Generative Adversarial Networks
by: Saraceno, Giovanni, et al.
Published: (2025)
by: Saraceno, Giovanni, et al.
Published: (2025)
MotionCFG: Boosting Motion Dynamics via Stochastic Concept Perturbation
by: Kim, Byungjun, et al.
Published: (2026)
by: Kim, Byungjun, et al.
Published: (2026)
TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation
by: Grimal, Paul, et al.
Published: (2023)
by: Grimal, Paul, et al.
Published: (2023)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
Towards Learning Foundation Models for Heuristic Functions to Solve Pathfinding Problems
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
by: Raffel, Matthew, et al.
Published: (2024)
by: Raffel, Matthew, et al.
Published: (2024)
LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions
by: Agostinelli, Victor, et al.
Published: (2024)
by: Agostinelli, Victor, et al.
Published: (2024)
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
by: Raffel, Matthew, et al.
Published: (2025)
by: Raffel, Matthew, et al.
Published: (2025)
State Diversity Matters in Offline Behavior Distillation
by: Lei, Shiye, et al.
Published: (2025)
by: Lei, Shiye, et al.
Published: (2025)
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
by: Chaudhary, Gaurav, et al.
Published: (2025)
by: Chaudhary, Gaurav, et al.
Published: (2025)
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
by: Tomihari, Akiyoshi, et al.
Published: (2026)
by: Tomihari, Akiyoshi, et al.
Published: (2026)
Aspects of human memory and Large Language Models
by: Janik, Romuald A.
Published: (2023)
by: Janik, Romuald A.
Published: (2023)
Pairwise Markov Chains for Volatility Forecasting
by: Azeraf, Elie
Published: (2024)
by: Azeraf, Elie
Published: (2024)
Similar Items
-
WARM: On the Benefits of Weight Averaged Reward Models
by: Ramé, Alexandre, et al.
Published: (2024) -
WARP: On the Benefits of Weight Averaged Rewarded Policies
by: Ramé, Alexandre, et al.
Published: (2024) -
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024) -
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025) -
Learning in Mean Field Games: A Survey
by: Laurière, Mathieu, et al.
Published: (2022)