Spectral Souping: A Unified Framework for Online Preference Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Chow, Yinlam, Tennenholtz, Guy, Yun, Ted, Harrison, James, Gretton, Arthur, Barreto, Andre, Dai, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spectral Bellman Method: Unifying Representation and Exploration in RL
by: Nabati, Ofir, et al.
Published: (2025)
by: Nabati, Ofir, et al.
Published: (2025)
DynaMITE-RL: A Dynamic Model for Improved Temporal Meta-Reinforcement Learning
by: Liang, Anthony, et al.
Published: (2024)
by: Liang, Anthony, et al.
Published: (2024)
Preference Adaptive and Sequential Text-to-Image Generation
by: Nabati, Ofir, et al.
Published: (2024)
by: Nabati, Ofir, et al.
Published: (2024)
Embedding-Aligned Language Models
by: Tennenholtz, Guy, et al.
Published: (2024)
by: Tennenholtz, Guy, et al.
Published: (2024)
Spectral Representation for Causal Estimation with Hidden Confounders
by: Sun, Haotian, et al.
Published: (2024)
by: Sun, Haotian, et al.
Published: (2024)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
by: Chow, Yinlam, et al.
Published: (2024)
by: Chow, Yinlam, et al.
Published: (2024)
Demystifying Embedding Spaces using Large Language Models
by: Tennenholtz, Guy, et al.
Published: (2023)
by: Tennenholtz, Guy, et al.
Published: (2023)
Perturbative methods for non-parametric instrumental variable
by: Bu, Wei, et al.
Published: (2026)
by: Bu, Wei, et al.
Published: (2026)
Semiparametric Efficient Test for Interpretable Distributional Treatment Effects
by: Zenati, Houssam, et al.
Published: (2026)
by: Zenati, Houssam, et al.
Published: (2026)
Kernel Single Proxy Control for Deterministic Confounding
by: Xu, Liyuan, et al.
Published: (2023)
by: Xu, Liyuan, et al.
Published: (2023)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
by: Chen, Zonghao, et al.
Published: (2025)
by: Chen, Zonghao, et al.
Published: (2025)
Diffusion Controller: Framework, Algorithms and Parameterization
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Optimal Rates for Vector-Valued Spectral Regularization Learning Algorithms
by: Meunier, Dimitri, et al.
Published: (2024)
by: Meunier, Dimitri, et al.
Published: (2024)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators
by: Linial, Ori, et al.
Published: (2024)
by: Linial, Ori, et al.
Published: (2024)
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)
by: Petrik, Marek, et al.
Published: (2023)
Demystifying Spectral Feature Learning for Instrumental Variable Regression
by: Meunier, Dimitri, et al.
Published: (2025)
by: Meunier, Dimitri, et al.
Published: (2025)
A Unified Data Representation Learning for Non-parametric Two-sample Testing
by: Tian, Xunye, et al.
Published: (2024)
by: Tian, Xunye, et al.
Published: (2024)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
by: Chegini, Atoosa, et al.
Published: (2024)
by: Chegini, Atoosa, et al.
Published: (2024)
Representation-Driven Reinforcement Learning
by: Nabati, Ofir, et al.
Published: (2023)
by: Nabati, Ofir, et al.
Published: (2023)
Regularized $f$-Divergence Kernel Tests
by: Ribero, Mónica, et al.
Published: (2026)
by: Ribero, Mónica, et al.
Published: (2026)
Doubly-Robust Estimation of Counterfactual Policy Mean Embeddings
by: Zenati, Houssam, et al.
Published: (2025)
by: Zenati, Houssam, et al.
Published: (2025)
Deep Proxy Causal Learning and its Application to Confounded Bandit Policy Evaluation
by: Xu, Liyuan, et al.
Published: (2021)
by: Xu, Liyuan, et al.
Published: (2021)
Descriptive History Representations: Learning Representations by Answering Questions
by: Tennenholtz, Guy, et al.
Published: (2025)
by: Tennenholtz, Guy, et al.
Published: (2025)
Interventional Processes for Causal Uncertainty Quantification
by: Dance, Hugh, et al.
Published: (2024)
by: Dance, Hugh, et al.
Published: (2024)
Kernel Treatment Effects with Adaptively Collected Data
by: Zenati, Houssam, et al.
Published: (2025)
by: Zenati, Houssam, et al.
Published: (2025)
A Distributional Analogue to the Successor Representation
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
Beyond RLHF: A Unified Theoretical Framework of Alignment
by: Yun, Jihun, et al.
Published: (2025)
by: Yun, Jihun, et al.
Published: (2025)
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
by: Sun, Shengyang, et al.
Published: (2025)
by: Sun, Shengyang, et al.
Published: (2025)
Near-Optimality of Contrastive Divergence Algorithms
by: Glaser, Pierre, et al.
Published: (2025)
by: Glaser, Pierre, et al.
Published: (2025)
Sequential Kernel Embedding for Mediated and Time-Varying Dose Response Curves
by: Singh, Rahul, et al.
Published: (2021)
by: Singh, Rahul, et al.
Published: (2021)
Self-Soupervision: Cooking Model Soups without Labels
by: Fuller, Anthony, et al.
Published: (2026)
by: Fuller, Anthony, et al.
Published: (2026)
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
by: Dai, Xiangxiang, et al.
Published: (2025)
by: Dai, Xiangxiang, et al.
Published: (2025)
Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression
by: Meunier, Dimitri, et al.
Published: (2025)
by: Meunier, Dimitri, et al.
Published: (2025)
Semiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity
by: Wornbard, Jakub, et al.
Published: (2026)
by: Wornbard, Jakub, et al.
Published: (2026)
Nonparametric Instrumental Regression via Kernel Methods is Minimax Optimal
by: Meunier, Dimitri, et al.
Published: (2024)
by: Meunier, Dimitri, et al.
Published: (2024)
Fast and Scalable Score-Based Kernel Calibration Tests
by: Glaser, Pierre, et al.
Published: (2025)
by: Glaser, Pierre, et al.
Published: (2025)
Towards Optimal Sobolev Norm Rates for the Vector-Valued Regularized Least-Squares Algorithm
by: Li, Zhu, et al.
Published: (2023)
by: Li, Zhu, et al.
Published: (2023)
Preference Alignment with Flow Matching
by: Kim, Minu, et al.
Published: (2024)
by: Kim, Minu, et al.
Published: (2024)
Deep MMD Gradient Flow without adversarial training
by: Galashov, Alexandre, et al.
Published: (2024)
by: Galashov, Alexandre, et al.
Published: (2024)
Similar Items
-
Spectral Bellman Method: Unifying Representation and Exploration in RL
by: Nabati, Ofir, et al.
Published: (2025) -
DynaMITE-RL: A Dynamic Model for Improved Temporal Meta-Reinforcement Learning
by: Liang, Anthony, et al.
Published: (2024) -
Preference Adaptive and Sequential Text-to-Image Generation
by: Nabati, Ofir, et al.
Published: (2024) -
Embedding-Aligned Language Models
by: Tennenholtz, Guy, et al.
Published: (2024) -
Spectral Representation for Causal Estimation with Hidden Confounders
by: Sun, Haotian, et al.
Published: (2024)