Knowledge Gradient for Preference Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Kaiwen, Gardner, Jacob R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Stochastic Natural Gradient Variational Inference
by: Wu, Kaiwen, et al.
Published: (2024)
by: Wu, Kaiwen, et al.
Published: (2024)
A Fast, Robust Elliptical Slice Sampling Implementation for Linearly Truncated Multivariate Normal Distributions
by: Wu, Kaiwen, et al.
Published: (2024)
by: Wu, Kaiwen, et al.
Published: (2024)
The Behavior and Convergence of Local Bayesian Optimization
by: Wu, Kaiwen, et al.
Published: (2023)
by: Wu, Kaiwen, et al.
Published: (2023)
Large-Scale Gaussian Processes via Alternating Projection
by: Wu, Kaiwen, et al.
Published: (2023)
by: Wu, Kaiwen, et al.
Published: (2023)
Computation-Aware Gaussian Processes: Model Selection And Linear-Time Inference
by: Wenger, Jonathan, et al.
Published: (2024)
by: Wenger, Jonathan, et al.
Published: (2024)
Demystifying SGD with Doubly Stochastic Gradients
by: Kim, Kyurae, et al.
Published: (2024)
by: Kim, Kyurae, et al.
Published: (2024)
On the Convergence of Black-Box Variational Inference
by: Kim, Kyurae, et al.
Published: (2023)
by: Kim, Kyurae, et al.
Published: (2023)
Reinforcement Learning via Value Gradient Flow
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Stochastic Gradient Variational Inference with Price's Gradient Estimator from Bures-Wasserstein to Parameter Space
by: Kim, Kyurae, et al.
Published: (2026)
by: Kim, Kyurae, et al.
Published: (2026)
Gradient Imbalance in Direct Preference Optimization
by: Ma, Qinwei, et al.
Published: (2025)
by: Ma, Qinwei, et al.
Published: (2025)
Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences
by: Liu, Ziang, et al.
Published: (2024)
by: Liu, Ziang, et al.
Published: (2024)
Collaborative Parameter Learning: Mitigating Forgetting via Parameter-Level Gradient Analysis
by: Yang, Mutian, et al.
Published: (2026)
by: Yang, Mutian, et al.
Published: (2026)
Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
by: Jones, Haydn, et al.
Published: (2026)
by: Jones, Haydn, et al.
Published: (2026)
Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
GI-SMN: Gradient Inversion Attack against Federated Learning without Prior Knowledge
by: Qian, Jin, et al.
Published: (2024)
by: Qian, Jin, et al.
Published: (2024)
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models
by: Mouiche, Inoussa
Published: (2026)
by: Mouiche, Inoussa
Published: (2026)
PMGDA: A Preference-based Multiple Gradient Descent Algorithm
by: Zhang, Xiaoyuan, et al.
Published: (2024)
by: Zhang, Xiaoyuan, et al.
Published: (2024)
Linear Convergence of Black-Box Variational Inference: Should We Stick the Landing?
by: Kim, Kyurae, et al.
Published: (2023)
by: Kim, Kyurae, et al.
Published: (2023)
Geometry-Aware Normalizing Wasserstein Flows for Optimal Causal Inference
by: Hou, Kaiwen
Published: (2023)
by: Hou, Kaiwen
Published: (2023)
FedSSP: Federated Graph Learning with Spectral Knowledge and Personalized Preference
by: Tan, Zihan, et al.
Published: (2024)
by: Tan, Zihan, et al.
Published: (2024)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
Balanced Gradient Sample Retrieval for Enhanced Knowledge Retention in Proxy-based Continual Learning
by: Xu, Hongye, et al.
Published: (2024)
by: Xu, Hongye, et al.
Published: (2024)
Preference-Guided Reinforcement Learning for Efficient Exploration
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Expert with Clustering: Hierarchical Online Preference Learning Framework
by: Zhou, Tianyue, et al.
Published: (2024)
by: Zhou, Tianyue, et al.
Published: (2024)
Tuning Sequential Monte Carlo Samplers via Greedy Incremental Divergence Minimization
by: Kim, Kyurae, et al.
Published: (2025)
by: Kim, Kyurae, et al.
Published: (2025)
Diffusion Classifier-Driven Reward for Offline Preference-based Reinforcement Learning
by: Pang, Teng, et al.
Published: (2025)
by: Pang, Teng, et al.
Published: (2025)
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
by: Ghosh, Udita, et al.
Published: (2025)
by: Ghosh, Udita, et al.
Published: (2025)
Contextual Preference Distribution Learning
by: Hudson, Benjamin, et al.
Published: (2026)
by: Hudson, Benjamin, et al.
Published: (2026)
MultiConfederated Learning: Inclusive Non-IID Data handling with Decentralized Federated Learning
by: Duchesne, Michael, et al.
Published: (2024)
by: Duchesne, Michael, et al.
Published: (2024)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
Preference-Based Gradient Estimation for ML-Guided Approximate Combinatorial Optimization
by: Mielke, Arman, et al.
Published: (2025)
by: Mielke, Arman, et al.
Published: (2025)
KG4Diagnosis: A Hierarchical Multi-Agent LLM Framework with Knowledge Graph Enhancement for Medical Diagnosis
by: Zuo, Kaiwen, et al.
Published: (2024)
by: Zuo, Kaiwen, et al.
Published: (2024)
Perseus: Leveraging Common Data Patterns with Curriculum Learning for More Robust Graph Neural Networks
by: Xia, Kaiwen, et al.
Published: (2024)
by: Xia, Kaiwen, et al.
Published: (2024)
Hindsight Preference Learning for Offline Preference-based Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2024)
by: Gao, Chen-Xiao, et al.
Published: (2024)
Provably Scalable Black-Box Variational Inference with Structured Variational Families
by: Ko, Joohwan, et al.
Published: (2024)
by: Ko, Joohwan, et al.
Published: (2024)
RankList -- A Listwise Preference Learning Framework for Predicting Subjective Preferences
by: Naini, Abinay Reddy, et al.
Published: (2025)
by: Naini, Abinay Reddy, et al.
Published: (2025)
Gradient Perturbation: Learning to Perturb Gradients for Adaptive Training
by: Li, Hua
Published: (2026)
by: Li, Hua
Published: (2026)
PrefPoE: Advantage-Guided Preference Fusion for Learning Where to Explore
by: Lin, Zhihao, et al.
Published: (2025)
by: Lin, Zhihao, et al.
Published: (2025)
The Central Role of the Loss Function in Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
Similar Items
-
Understanding Stochastic Natural Gradient Variational Inference
by: Wu, Kaiwen, et al.
Published: (2024) -
A Fast, Robust Elliptical Slice Sampling Implementation for Linearly Truncated Multivariate Normal Distributions
by: Wu, Kaiwen, et al.
Published: (2024) -
The Behavior and Convergence of Local Bayesian Optimization
by: Wu, Kaiwen, et al.
Published: (2023) -
Large-Scale Gaussian Processes via Alternating Projection
by: Wu, Kaiwen, et al.
Published: (2023) -
Computation-Aware Gaussian Processes: Model Selection And Linear-Time Inference
by: Wenger, Jonathan, et al.
Published: (2024)