Diffusion Alignment Beyond KL: Variance Minimisation as Effective Policy Optimiser
Fuente:
arXiv
Saved in:
| Main Authors: | Ou, Zijing, Si, Jacob, Zhu, Junyi, Bohdal, Ondrej, Ozay, Mete, Ceritli, Taha, Li, Yingzhen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Clustering-driven Memory Compression for On-device Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2026)
by: Bohdal, Ondrej, et al.
Published: (2026)
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
by: Shenaj, Donald, et al.
Published: (2025)
by: Shenaj, Donald, et al.
Published: (2025)
MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
by: Ceritli, Taha, et al.
Published: (2025)
by: Ceritli, Taha, et al.
Published: (2025)
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
by: Hosseini, Peyman, et al.
Published: (2025)
by: Hosseini, Peyman, et al.
Published: (2025)
Data-driven Clustering and Merging of Adapters for On-device Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2026)
by: Bohdal, Ondrej, et al.
Published: (2026)
TabRep: Training Tabular Diffusion Models with a Simple and Effective Continuous Representation
by: Si, Jacob, et al.
Published: (2025)
by: Si, Jacob, et al.
Published: (2025)
LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
by: Shenaj, Donald, et al.
Published: (2024)
by: Shenaj, Donald, et al.
Published: (2024)
Inference-Time Scaling of Discrete Diffusion Models via Importance Weighting and Optimal Proposal Design
by: Ou, Zijing, et al.
Published: (2025)
by: Ou, Zijing, et al.
Published: (2025)
Guided Model Merging for Hybrid Data Learning: Leveraging Centralized Data to Refine Decentralized Models
by: Zhu, Junyi, et al.
Published: (2025)
by: Zhu, Junyi, et al.
Published: (2025)
Efficient Compositional Multi-tasking for On-device Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2025)
by: Bohdal, Ondrej, et al.
Published: (2025)
Discrete Neural Flow Samplers with Locally Equivariant Transformer
by: Ou, Zijing, et al.
Published: (2025)
by: Ou, Zijing, et al.
Published: (2025)
Neural Flow Samplers with Shortcut Models
by: Chen, Wuhao, et al.
Published: (2025)
by: Chen, Wuhao, et al.
Published: (2025)
Randomized Asymmetric Chain of LoRA: The First Meaningful Theoretical Framework for Low-Rank Adaptation
by: Malinovsky, Grigory, et al.
Published: (2024)
by: Malinovsky, Grigory, et al.
Published: (2024)
Mutual Information Multinomial Estimation
by: Chen, Yanzhi, et al.
Published: (2024)
by: Chen, Yanzhi, et al.
Published: (2024)
Energy-Based Modelling for Discrete and Mixed Data via Heat Equations on Structured Spaces
by: Schröder, Tobias, et al.
Published: (2024)
by: Schröder, Tobias, et al.
Published: (2024)
Improving Probabilistic Diffusion Models With Optimal Diagonal Covariance Matching
by: Ou, Zijing, et al.
Published: (2024)
by: Ou, Zijing, et al.
Published: (2024)
VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning
by: Zong, Yongshuo, et al.
Published: (2024)
by: Zong, Yongshuo, et al.
Published: (2024)
HOP to the Next Tasks and Domains for Continual Learning in NLP
by: Michieli, Umberto, et al.
Published: (2024)
by: Michieli, Umberto, et al.
Published: (2024)
On the Limitations of General Purpose Domain Generalisation Methods
by: Gouk, Henry, et al.
Published: (2022)
by: Gouk, Henry, et al.
Published: (2022)
Feed-Forward Latent Domain Adaptation
by: Bohdal, Ondrej, et al.
Published: (2022)
by: Bohdal, Ondrej, et al.
Published: (2022)
Navigating Noise: A Study of How Noise Influences Generalisation and Calibration of Neural Networks
by: Ferianc, Martin, et al.
Published: (2023)
by: Ferianc, Martin, et al.
Published: (2023)
HiBBO: HiPPO-based Space Consistency for High-dimensional Bayesian Optimisation
by: Xuan, Junyu, et al.
Published: (2025)
by: Xuan, Junyu, et al.
Published: (2025)
Memorized Images in Diffusion Models share a Subspace that can be Located and Deleted
by: Chavhan, Ruchika, et al.
Published: (2024)
by: Chavhan, Ruchika, et al.
Published: (2024)
Memory-Driven Self-Improvement for Decision Making with Large Language Models
by: Yan, Xue, et al.
Published: (2025)
by: Yan, Xue, et al.
Published: (2025)
Projective Proximal Gradient Descent for A Class of Nonconvex Nonsmooth Optimization Problems: Fast Convergence Without Kurdyka-Lojasiewicz (KL) Property
by: Yang, Yingzhen, et al.
Published: (2023)
by: Yang, Yingzhen, et al.
Published: (2023)
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
by: Camuffo, Elena, et al.
Published: (2025)
by: Camuffo, Elena, et al.
Published: (2025)
VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation
by: Wang, Leyang, et al.
Published: (2025)
by: Wang, Leyang, et al.
Published: (2025)
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
by: Zong, Yongshuo, et al.
Published: (2024)
by: Zong, Yongshuo, et al.
Published: (2024)
Tangent Space Fine-Tuning for Directional Preference Alignment in Large Language Models
by: Erdogan, Mete
Published: (2026)
by: Erdogan, Mete
Published: (2026)
On-device System of Compositional Multi-tasking in Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2025)
by: Bohdal, Ondrej, et al.
Published: (2025)
Extraction and Recovery of Spatio-Temporal Structure in Latent Dynamics Alignment with Diffusion Models
by: Wang, Yule, et al.
Published: (2023)
by: Wang, Yule, et al.
Published: (2023)
Feature-Space Generative Models for One-Shot Class-Incremental Learning
by: Foster, Jack, et al.
Published: (2026)
by: Foster, Jack, et al.
Published: (2026)
Forward KL Regularized Preference Optimization for Aligning Diffusion Policies
by: Shan, Zhao, et al.
Published: (2024)
by: Shan, Zhao, et al.
Published: (2024)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Understanding NTK Variance in Implicit Neural Representations
by: Ou, Chengguang, et al.
Published: (2025)
by: Ou, Chengguang, et al.
Published: (2025)
Differentially Private Clustered Federated Learning with Privacy-Preserving Initialization and Normality-Driven Aggregation
by: Xu, Jie, et al.
Published: (2026)
by: Xu, Jie, et al.
Published: (2026)
Diffusion on Graph: Augmentation of Graph Structure for Node Classification
by: Wang, Yancheng, et al.
Published: (2025)
by: Wang, Yancheng, et al.
Published: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
Variational Uncertainty Decomposition for In-Context Learning
by: Jayasekera, I. Shavindra, et al.
Published: (2025)
by: Jayasekera, I. Shavindra, et al.
Published: (2025)
Similar Items
-
Clustering-driven Memory Compression for On-device Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2026) -
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
by: Shenaj, Donald, et al.
Published: (2025) -
MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
by: Bini, Massimo, et al.
Published: (2025) -
HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
by: Ceritli, Taha, et al.
Published: (2025) -
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
by: Hosseini, Peyman, et al.
Published: (2025)