Gespeichert in:
| Hauptverfasser: | Zhao, Yao, Jun, Kwang-Sung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.01581 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Coverage Improvement and Fast Convergence of On-policy Preference Learning
von: Kim, Juno, et al.
Veröffentlicht: (2026)
von: Kim, Juno, et al.
Veröffentlicht: (2026)
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
von: Jun, Kwang-Sung, et al.
Veröffentlicht: (2024)
von: Jun, Kwang-Sung, et al.
Veröffentlicht: (2024)
GL-LowPopArt: A Nearly Instance-Wise Minimax-Optimal Estimator for Generalized Low-Rank Trace Regression
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search
von: Nguyen, Tuan Ngo, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan Ngo, et al.
Veröffentlicht: (2024)
Adaptive Experimentation When You Can't Experiment
von: Zhao, Yao, et al.
Veröffentlicht: (2024)
von: Zhao, Yao, et al.
Veröffentlicht: (2024)
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
von: Lee, Junghyun, et al.
Veröffentlicht: (2024)
von: Lee, Junghyun, et al.
Veröffentlicht: (2024)
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
von: Li, Yinan, et al.
Veröffentlicht: (2025)
von: Li, Yinan, et al.
Veröffentlicht: (2025)
Regularized Online RLHF with Generalized Bilinear Preferences
von: Lee, Junghyun, et al.
Veröffentlicht: (2026)
von: Lee, Junghyun, et al.
Veröffentlicht: (2026)
Adversarial Preference Learning for Robust LLM Alignment
von: Wang, Yuanfu, et al.
Veröffentlicht: (2025)
von: Wang, Yuanfu, et al.
Veröffentlicht: (2025)
Data Selection for LLM Alignment Using Fine-Grained Preferences
von: Zhang, Jia, et al.
Veröffentlicht: (2025)
von: Zhang, Jia, et al.
Veröffentlicht: (2025)
Beyond RLHF: A Unified Theoretical Framework of Alignment
von: Yun, Jihun, et al.
Veröffentlicht: (2025)
von: Yun, Jihun, et al.
Veröffentlicht: (2025)
$\varepsilon$-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
von: Li, Yinan, et al.
Veröffentlicht: (2026)
von: Li, Yinan, et al.
Veröffentlicht: (2026)
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
von: Jang, Kyoungseok, et al.
Veröffentlicht: (2024)
von: Jang, Kyoungseok, et al.
Veröffentlicht: (2024)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
von: Qin, Hao, et al.
Veröffentlicht: (2023)
von: Qin, Hao, et al.
Veröffentlicht: (2023)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
Distributional Preference Alignment of LLMs via Optimal Transport
von: Melnyk, Igor, et al.
Veröffentlicht: (2024)
von: Melnyk, Igor, et al.
Veröffentlicht: (2024)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
QuietPaw: Learning Quadrupedal Locomotion with Versatile Noise Preference Alignment
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
Can Revealed Preferences Clarify LLM Alignment and Steering?
von: Yamin, Khurram, et al.
Veröffentlicht: (2026)
von: Yamin, Khurram, et al.
Veröffentlicht: (2026)
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
von: Qin, Hao, et al.
Veröffentlicht: (2025)
von: Qin, Hao, et al.
Veröffentlicht: (2025)
Sample Efficient Preference Alignment in LLMs via Active Exploration
von: Mehta, Viraj, et al.
Veröffentlicht: (2023)
von: Mehta, Viraj, et al.
Veröffentlicht: (2023)
Active Query Synthesis for Preference Learning
von: Nadagouda, Namrata, et al.
Veröffentlicht: (2026)
von: Nadagouda, Namrata, et al.
Veröffentlicht: (2026)
Active Learning for Direct Preference Optimization
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
Alignment with Preference Optimization Is All You Need for LLM Safety
von: Alami, Reda, et al.
Veröffentlicht: (2024)
von: Alami, Reda, et al.
Veröffentlicht: (2024)
Holistic Utility Preference Learning for Listwise Alignment
von: Zhou, Jiacong, et al.
Veröffentlicht: (2024)
von: Zhou, Jiacong, et al.
Veröffentlicht: (2024)
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2026)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2026)
Learning Explainable Dense Reward Shapes via Bayesian Optimization
von: Koo, Ryan, et al.
Veröffentlicht: (2025)
von: Koo, Ryan, et al.
Veröffentlicht: (2025)
RANA: Robust Active Learning for Noisy Network Alignment
von: Nan, Yixuan, et al.
Veröffentlicht: (2025)
von: Nan, Yixuan, et al.
Veröffentlicht: (2025)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
Better-than-KL PAC-Bayes Bounds
von: Kuzborskij, Ilja, et al.
Veröffentlicht: (2024)
von: Kuzborskij, Ilja, et al.
Veröffentlicht: (2024)
Offline Clustering of Preference Learning with Active-data Augmentation
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
Active Preference Learning for Ordering Items In- and Out-of-sample
von: Bergström, Herman, et al.
Veröffentlicht: (2024)
von: Bergström, Herman, et al.
Veröffentlicht: (2024)
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
von: Chen, Haoxian, et al.
Veröffentlicht: (2024)
von: Chen, Haoxian, et al.
Veröffentlicht: (2024)
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
Preference Alignment with Flow Matching
von: Kim, Minu, et al.
Veröffentlicht: (2024)
von: Kim, Minu, et al.
Veröffentlicht: (2024)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
von: Xu, Zaiyan, et al.
Veröffentlicht: (2025)
von: Xu, Zaiyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Coverage Improvement and Fast Convergence of On-policy Preference Learning
von: Kim, Juno, et al.
Veröffentlicht: (2026) -
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
von: Jun, Kwang-Sung, et al.
Veröffentlicht: (2024) -
GL-LowPopArt: A Nearly Instance-Wise Minimax-Optimal Estimator for Generalized Low-Rank Trace Regression
von: Lee, Junghyun, et al.
Veröffentlicht: (2025) -
HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search
von: Nguyen, Tuan Ngo, et al.
Veröffentlicht: (2024) -
Adaptive Experimentation When You Can't Experiment
von: Zhao, Yao, et al.
Veröffentlicht: (2024)