Loss Functions and Operators Generated by f-Divergences
Fuente:
arXiv
Saved in:
| Main Authors: | Roulet, Vincent, Liu, Tianlin, Vieillard, Nino, Sander, Michael E., Blondel, Mathieu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint Learning of Energy-based Models and their Partition Function
by: Sander, Michael E., et al.
Published: (2025)
by: Sander, Michael E., et al.
Published: (2025)
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
by: Blondel, Mathieu, et al.
Published: (2025)
by: Blondel, Mathieu, et al.
Published: (2025)
The Elements of Differentiable Programming
by: Blondel, Mathieu, et al.
Published: (2024)
by: Blondel, Mathieu, et al.
Published: (2024)
Differentiable Knapsack and Top-k Operators via Dynamic Programming
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Routers in Vision Mixture of Experts: An Empirical Study
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024)
by: Roulet, Vincent, et al.
Published: (2024)
How do Transformers perform In-Context Autoregressive Learning?
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Learning with Fitzpatrick Losses
by: Rakotomandimby, Seta, et al.
Published: (2024)
by: Rakotomandimby, Seta, et al.
Published: (2024)
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
On Weak-to-Strong Generalization and f-Divergence
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
Empirical Risk Minimization with $f$-Divergence Regularization
by: Daunas, Francisco, et al.
Published: (2026)
by: Daunas, Francisco, et al.
Published: (2026)
Learning with Local Search MCMC Layers
by: Vivier-Ardisson, Germain, et al.
Published: (2025)
by: Vivier-Ardisson, Germain, et al.
Published: (2025)
Partition Function Estimation under Bounded f-Divergence
by: Block, Adam, et al.
Published: (2026)
by: Block, Adam, et al.
Published: (2026)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
by: Agarwal, Rishabh, et al.
Published: (2023)
by: Agarwal, Rishabh, et al.
Published: (2023)
Variational f-divergence Minimization
by: Zhang, Mingtian, et al.
Published: (2019)
by: Zhang, Mingtian, et al.
Published: (2019)
Regularized Large Neighborhood Search
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
Equivalence of the Empirical Risk Minimization to Regularization on the Family of f-Divergences
by: Daunas, Francisco, et al.
Published: (2024)
by: Daunas, Francisco, et al.
Published: (2024)
Generalization Error of $f$-Divergence Stabilized Algorithms via Duality
by: Daunas, Francisco, et al.
Published: (2025)
by: Daunas, Francisco, et al.
Published: (2025)
Minimizing $f$-Divergences by Interpolating Velocity Fields
by: Liu, Song, et al.
Published: (2023)
by: Liu, Song, et al.
Published: (2023)
Jensen-Shannon Divergence Based Novel Loss Functions for Bayesian Neural Networks
by: Thiagarajan, Ponkrshnan, et al.
Published: (2022)
by: Thiagarajan, Ponkrshnan, et al.
Published: (2022)
Regularized $f$-Divergence Kernel Tests
by: Ribero, Mónica, et al.
Published: (2026)
by: Ribero, Mónica, et al.
Published: (2026)
Decoding-time Realignment of Language Models
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
Generalized Kullback-Leibler Divergence Loss
by: Cui, Jiequan, et al.
Published: (2025)
by: Cui, Jiequan, et al.
Published: (2025)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023)
by: Roulet, Vincent, et al.
Published: (2023)
Sample Compression Unleashed: New Generalization Bounds for Real Valued Losses
by: Bazinet, Mathieu, et al.
Published: (2024)
by: Bazinet, Mathieu, et al.
Published: (2024)
WARM: On the Benefits of Weight Averaged Reward Models
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment
by: Haldar, Rajdeep, et al.
Published: (2026)
by: Haldar, Rajdeep, et al.
Published: (2026)
MVG-CRPS: A Robust Loss Function for Multivariate Probabilistic Forecasting
by: Zheng, Vincent Zhihao, et al.
Published: (2024)
by: Zheng, Vincent Zhihao, et al.
Published: (2024)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
by: Pipano, Idan, et al.
Published: (2026)
by: Pipano, Idan, et al.
Published: (2026)
Robust Semi-supervised Learning via $f$-Divergence and $α$-Rényi Divergence
by: Aminian, Gholamali, et al.
Published: (2024)
by: Aminian, Gholamali, et al.
Published: (2024)
Project and Generate: Divergence-Free Neural Operators for Incompressible Flows
by: Li, Xigui, et al.
Published: (2026)
by: Li, Xigui, et al.
Published: (2026)
Clustering in Deep Stochastic Transformers
by: Fedorov, Lev, et al.
Published: (2026)
by: Fedorov, Lev, et al.
Published: (2026)
Bounding Neyman-Pearson Region with $f$-Divergences
by: Mullhaupt, Andrew, et al.
Published: (2025)
by: Mullhaupt, Andrew, et al.
Published: (2025)
Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing
by: Xiao, Yisong, et al.
Published: (2023)
by: Xiao, Yisong, et al.
Published: (2023)
An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence
by: Zhang, Qizhen, et al.
Published: (2026)
by: Zhang, Qizhen, et al.
Published: (2026)
Regularization via f-Divergence: An Application to Multi-Oxide Spectroscopic Analysis
by: Li, Weizhi, et al.
Published: (2025)
by: Li, Weizhi, et al.
Published: (2025)
Decoupled Kullback-Leibler Divergence Loss
by: Cui, Jiequan, et al.
Published: (2023)
by: Cui, Jiequan, et al.
Published: (2023)
Relationship between Hölder Divergence and Functional Density Power Divergence: Intersection and Generalization
by: Kobayashi, Masahiro
Published: (2025)
by: Kobayashi, Masahiro
Published: (2025)
Robust Offline Reinforcement Learning with Linearly Structured f-Divergence Regularization
by: Tang, Cheng, et al.
Published: (2024)
by: Tang, Cheng, et al.
Published: (2024)
Similar Items
-
Joint Learning of Energy-based Models and their Partition Function
by: Sander, Michael E., et al.
Published: (2025) -
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
by: Blondel, Mathieu, et al.
Published: (2025) -
The Elements of Differentiable Programming
by: Blondel, Mathieu, et al.
Published: (2024) -
Differentiable Knapsack and Top-k Operators via Dynamic Programming
by: Vivier-Ardisson, Germain, et al.
Published: (2026) -
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)