Nonparametric Bayesian Optimization for General Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zishi, Ren, Tao, Peng, Yijie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sample-Efficient "Clustering and Conquer" Procedures for Parallel Large-Scale Ranking and Selection
von: Zhang, Zishi, et al.
Veröffentlicht: (2024)
von: Zhang, Zishi, et al.
Veröffentlicht: (2024)
Optimal low-rank stochastic gradient estimation for LLM training
von: Li, Zehao, et al.
Veröffentlicht: (2026)
von: Li, Zehao, et al.
Veröffentlicht: (2026)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
von: Ren, Tao, et al.
Veröffentlicht: (2025)
von: Ren, Tao, et al.
Veröffentlicht: (2025)
Cold-Start Forecasting of New Product Life-Cycles via Conditional Diffusion Models
von: Zhou, Ruihan, et al.
Veröffentlicht: (2026)
von: Zhou, Ruihan, et al.
Veröffentlicht: (2026)
FLOPS: Forward Learning with OPtimal Sampling
von: Ren, Tao, et al.
Veröffentlicht: (2024)
von: Ren, Tao, et al.
Veröffentlicht: (2024)
Omni-Masked Gradient Descent: Memory-Efficient Optimization via Mask Traversal with Improved Convergence
von: Yang, Hui, et al.
Veröffentlicht: (2026)
von: Yang, Hui, et al.
Veröffentlicht: (2026)
Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer
von: Ren, Tao, et al.
Veröffentlicht: (2025)
von: Ren, Tao, et al.
Veröffentlicht: (2025)
Bayesian Nonparametrics Meets Data-Driven Distributionally Robust Optimization
von: Bariletto, Nicola, et al.
Veröffentlicht: (2024)
von: Bariletto, Nicola, et al.
Veröffentlicht: (2024)
Bayesian Nonparametrics: An Alternative to Deep Learning
von: Moraffah, Bahman
Veröffentlicht: (2024)
von: Moraffah, Bahman
Veröffentlicht: (2024)
Learning Explainable Dense Reward Shapes via Bayesian Optimization
von: Koo, Ryan, et al.
Veröffentlicht: (2025)
von: Koo, Ryan, et al.
Veröffentlicht: (2025)
BN-Pool: Bayesian Nonparametric Pooling for Graphs
von: Castellana, Daniele, et al.
Veröffentlicht: (2025)
von: Castellana, Daniele, et al.
Veröffentlicht: (2025)
Bayesian Nonparametric Mixed-Effect ODEs with Gaussian Processes
von: Martinelli, Julien, et al.
Veröffentlicht: (2026)
von: Martinelli, Julien, et al.
Veröffentlicht: (2026)
On Cold Posteriors of Probabilistic Neural Networks: Understanding the Cold Posterior Effect and A New Way to Learn Cold Posteriors with Tight Generalization Guarantees
von: Zhang, Yijie
Veröffentlicht: (2024)
von: Zhang, Yijie
Veröffentlicht: (2024)
Nonparametric Teaching for Graph Property Learners
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
Optimal Initialization of Batch Bayesian Optimization
von: Ren, Jiuge, et al.
Veröffentlicht: (2024)
von: Ren, Jiuge, et al.
Veröffentlicht: (2024)
A Deep Bayesian Nonparametric Framework for Robust Mutual Information Estimation
von: Fazeliasl, Forough, et al.
Veröffentlicht: (2025)
von: Fazeliasl, Forough, et al.
Veröffentlicht: (2025)
A Bayesian Nonparametric Perspective on Mahalanobis Distance for Out of Distribution Detection
von: Linderman, Randolph W., et al.
Veröffentlicht: (2025)
von: Linderman, Randolph W., et al.
Veröffentlicht: (2025)
Generative Modeling with Multi-Instance Reward Learning for E-commerce Creative Optimization
von: Gu, Qiaolei, et al.
Veröffentlicht: (2025)
von: Gu, Qiaolei, et al.
Veröffentlicht: (2025)
Bayesian Reward Models for LLM Alignment
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
Efficient Nonparametric Tensor Decomposition for Binary and Count Data
von: Tao, Zerui, et al.
Veröffentlicht: (2024)
von: Tao, Zerui, et al.
Veröffentlicht: (2024)
A New Stochastic Approximation Method for Gradient-based Simulated Parameter Estimation
von: Li, Zehao, et al.
Veröffentlicht: (2025)
von: Li, Zehao, et al.
Veröffentlicht: (2025)
Heterogeneous Ordinal Structure Learning with Bayesian Nonparametric Complexity Discovery
von: Rafe, Amir, et al.
Veröffentlicht: (2026)
von: Rafe, Amir, et al.
Veröffentlicht: (2026)
SPDE Methods for Nonparametric Bayesian Posterior Contraction and Laplace Approximation
von: Alberola-Boloix, Enric, et al.
Veröffentlicht: (2026)
von: Alberola-Boloix, Enric, et al.
Veröffentlicht: (2026)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
Sharper Generalization Bounds for Transformer
von: Li, Yawen, et al.
Veröffentlicht: (2026)
von: Li, Yawen, et al.
Veröffentlicht: (2026)
Data-Driven DRO and Economic Decision Theory: An Analytical Synthesis With Bayesian Nonparametric Advancements
von: Bariletto, Nicola, et al.
Veröffentlicht: (2024)
von: Bariletto, Nicola, et al.
Veröffentlicht: (2024)
T-REG: Preference Optimization with Token-Level Reward Regularization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Stochastic Approximation Methods for Distortion Risk Measure Optimization
von: Jiang, Jinyang, et al.
Veröffentlicht: (2025)
von: Jiang, Jinyang, et al.
Veröffentlicht: (2025)
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
von: Cho, Brian, et al.
Veröffentlicht: (2024)
von: Cho, Brian, et al.
Veröffentlicht: (2024)
Likelihood-Free Adaptive Bayesian Inference via Nonparametric Distribution Matching
von: Lu, Wenhui Sophia, et al.
Veröffentlicht: (2025)
von: Lu, Wenhui Sophia, et al.
Veröffentlicht: (2025)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
Causal Representation Learning from General Environments under Nonparametric Mixing
von: Ng, Ignavier, et al.
Veröffentlicht: (2026)
von: Ng, Ignavier, et al.
Veröffentlicht: (2026)
Trajectory-Oriented Policy Optimization with Sparse Rewards
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Causal Bayesian Optimization via Exogenous Distribution Learning
von: Ren, Shaogang, et al.
Veröffentlicht: (2024)
von: Ren, Shaogang, et al.
Veröffentlicht: (2024)
Direct Regret Optimization in Bayesian Optimization
von: Zhang, Fengxue, et al.
Veröffentlicht: (2025)
von: Zhang, Fengxue, et al.
Veröffentlicht: (2025)
Generative Bayesian Optimization: Generative Models as Acquisition Functions
von: Oliveira, Rafael, et al.
Veröffentlicht: (2025)
von: Oliveira, Rafael, et al.
Veröffentlicht: (2025)
Bayesian Inverse Reinforcement Learning for Non-Markovian Rewards
von: Topper, Noah, et al.
Veröffentlicht: (2024)
von: Topper, Noah, et al.
Veröffentlicht: (2024)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Sample-Efficient "Clustering and Conquer" Procedures for Parallel Large-Scale Ranking and Selection
von: Zhang, Zishi, et al.
Veröffentlicht: (2024) -
Optimal low-rank stochastic gradient estimation for LLM training
von: Li, Zehao, et al.
Veröffentlicht: (2026) -
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
von: Zhang, Zishi, et al.
Veröffentlicht: (2026) -
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
von: Ren, Tao, et al.
Veröffentlicht: (2025) -
Cold-Start Forecasting of New Product Life-Cycles via Conditional Diffusion Models
von: Zhou, Ruihan, et al.
Veröffentlicht: (2026)