Soft Preference Optimization: Aligning Language Models to Expert Distributions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sharifnassab, Arsalan, Salehkaleybar, Saber, Ghiassian, Sina, Kanoria, Surya, Schuurmans, Dale |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2024)
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2024)
Order Optimal Bounds for One-Shot Federated Learning over non-Convex Loss Functions
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2021)
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2021)
Step-size Optimization for Continual Learning
von: Degris, Thomas, et al.
Veröffentlicht: (2024)
von: Degris, Thomas, et al.
Veröffentlicht: (2024)
Causal Effect Identification in Heterogeneous Environments from Higher-Order Moments
von: Kivva, Yaroslav, et al.
Veröffentlicht: (2025)
von: Kivva, Yaroslav, et al.
Veröffentlicht: (2025)
Near-Optimal Experiment Design in Linear non-Gaussian Cyclic Models
von: Sharifian, Ehsan, et al.
Veröffentlicht: (2025)
von: Sharifian, Ehsan, et al.
Veröffentlicht: (2025)
Multi-Domain Causal Discovery in Bijective Causal Models
von: Jalaldoust, Kasra, et al.
Veröffentlicht: (2025)
von: Jalaldoust, Kasra, et al.
Veröffentlicht: (2025)
In-context Exploration-Exploitation for Reinforcement Learning
von: Dai, Zhenwen, et al.
Veröffentlicht: (2024)
von: Dai, Zhenwen, et al.
Veröffentlicht: (2024)
Learning in complex action spaces without policy gradients
von: Tavakoli, Arash, et al.
Veröffentlicht: (2024)
von: Tavakoli, Arash, et al.
Veröffentlicht: (2024)
Learning Unknown Intervention Targets in Structural Causal Models from Heterogeneous Data
von: Yang, Yuqin, et al.
Veröffentlicht: (2023)
von: Yang, Yuqin, et al.
Veröffentlicht: (2023)
ACTIVA: Amortized Causal Effect Estimation via Transformer-based Variational Autoencoder
von: Sauter, Andreas, et al.
Veröffentlicht: (2025)
von: Sauter, Andreas, et al.
Veröffentlicht: (2025)
Intentional Updates for Streaming Reinforcement Learning
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2026)
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2026)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
Aligning Diffusion Language Models via Unpaired Preference Optimization
von: Jindal, Vaibhav, et al.
Veröffentlicht: (2025)
von: Jindal, Vaibhav, et al.
Veröffentlicht: (2025)
Auxiliary task discovery through generate-and-test
von: Rafiee, Banafsheh, et al.
Veröffentlicht: (2022)
von: Rafiee, Banafsheh, et al.
Veröffentlicht: (2022)
Deep Reinforcement Learning for Inventory Networks: Toward Reliable Policy Optimization
von: Alvo, Matias, et al.
Veröffentlicht: (2023)
von: Alvo, Matias, et al.
Veröffentlicht: (2023)
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
von: Nguyen, Thanh Thi, et al.
Veröffentlicht: (2025)
von: Nguyen, Thanh Thi, et al.
Veröffentlicht: (2025)
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
von: Alvo, Matias, et al.
Veröffentlicht: (2026)
von: Alvo, Matias, et al.
Veröffentlicht: (2026)
SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
Aligning CodeLLMs with Direct Preference Optimization
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
Spectral Representation-based Reinforcement Learning
von: Gao, Chenxiao, et al.
Veröffentlicht: (2025)
von: Gao, Chenxiao, et al.
Veröffentlicht: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
von: Zhang, Hongming, et al.
Veröffentlicht: (2023)
von: Zhang, Hongming, et al.
Veröffentlicht: (2023)
Geometric-Averaged Preference Optimization for Soft Preference Labels
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
von: Li, Junsong, et al.
Veröffentlicht: (2025)
von: Li, Junsong, et al.
Veröffentlicht: (2025)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Causal Reasoning in Pieces: Modular In-Context Learning for Causal Discovery
von: Kadziolka, Kacper, et al.
Veröffentlicht: (2025)
von: Kadziolka, Kacper, et al.
Veröffentlicht: (2025)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
von: Liu, Yinhong, et al.
Veröffentlicht: (2024)
von: Liu, Yinhong, et al.
Veröffentlicht: (2024)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
von: Rastegar, Reza
Veröffentlicht: (2026)
von: Rastegar, Reza
Veröffentlicht: (2026)
From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
von: Zhu, Mingkang, et al.
Veröffentlicht: (2025)
von: Zhu, Mingkang, et al.
Veröffentlicht: (2025)
Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph
von: Liu, Ning, et al.
Veröffentlicht: (2026)
von: Liu, Ning, et al.
Veröffentlicht: (2026)
Scaffold-Conditioned Preference Triplets for Controllable Molecular Optimization with Large Language Models
von: Xiong, Yi, et al.
Veröffentlicht: (2026)
von: Xiong, Yi, et al.
Veröffentlicht: (2026)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
von: Qin, Kai, et al.
Veröffentlicht: (2025)
von: Qin, Kai, et al.
Veröffentlicht: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
von: Xu, Zaiyan, et al.
Veröffentlicht: (2025)
von: Xu, Zaiyan, et al.
Veröffentlicht: (2025)
Scalable Diffusion for Materials Generation
von: Yang, Sherry, et al.
Veröffentlicht: (2023)
von: Yang, Sherry, et al.
Veröffentlicht: (2023)
Structure-Aligned Protein Language Model
von: Chen, Can, et al.
Veröffentlicht: (2025)
von: Chen, Can, et al.
Veröffentlicht: (2025)
ROPO: Robust Preference Optimization for Large Language Models
von: Liang, Xize, et al.
Veröffentlicht: (2024)
von: Liang, Xize, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2024) -
Order Optimal Bounds for One-Shot Federated Learning over non-Convex Loss Functions
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2021) -
Step-size Optimization for Continual Learning
von: Degris, Thomas, et al.
Veröffentlicht: (2024) -
Causal Effect Identification in Heterogeneous Environments from Higher-Order Moments
von: Kivva, Yaroslav, et al.
Veröffentlicht: (2025) -
Near-Optimal Experiment Design in Linear non-Gaussian Cyclic Models
von: Sharifian, Ehsan, et al.
Veröffentlicht: (2025)