Saved in:
| Main Authors: | Kim, Juno, Yun, Jihun, Lee, Jason D., Jun, Kwang-Sung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.08421 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond RLHF: A Unified Theoretical Framework of Alignment
by: Yun, Jihun, et al.
Published: (2025)
by: Yun, Jihun, et al.
Published: (2025)
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
by: Zhao, Yao, et al.
Published: (2026)
by: Zhao, Yao, et al.
Published: (2026)
Regularized Online RLHF with Generalized Bilinear Preferences
by: Lee, Junghyun, et al.
Published: (2026)
by: Lee, Junghyun, et al.
Published: (2026)
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
by: Lee, Junghyun, et al.
Published: (2024)
by: Lee, Junghyun, et al.
Published: (2024)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
by: Lee, Junghyun, et al.
Published: (2023)
by: Lee, Junghyun, et al.
Published: (2023)
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
by: Jun, Kwang-Sung, et al.
Published: (2024)
by: Jun, Kwang-Sung, et al.
Published: (2024)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026)
by: Kim, Juno, et al.
Published: (2026)
XB-MAML: Learning Expandable Basis Parameters for Effective Meta-Learning with Wide Task Coverage
by: Lee, Jae-Jun, et al.
Published: (2024)
by: Lee, Jae-Jun, et al.
Published: (2024)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Uniform Spectral Growth and Convergence of Muon in LoRA-Style Matrix Factorization
by: Kang, Changmin, et al.
Published: (2026)
by: Kang, Changmin, et al.
Published: (2026)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Elucidating Subspace Perturbation in Zeroth-Order Optimization: Theory and Practice at Scale
by: Park, Sihwan, et al.
Published: (2025)
by: Park, Sihwan, et al.
Published: (2025)
SeedFlood: A Step Toward Scalable Decentralized Training of LLMs
by: Kim, Jihun, et al.
Published: (2026)
by: Kim, Jihun, et al.
Published: (2026)
GL-LowPopArt: A Nearly Instance-Wise Minimax-Optimal Estimator for Generalized Low-Rank Trace Regression
by: Lee, Junghyun, et al.
Published: (2025)
by: Lee, Junghyun, et al.
Published: (2025)
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
by: Balagopalan, Kapilan, et al.
Published: (2024)
by: Balagopalan, Kapilan, et al.
Published: (2024)
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
by: Li, Yinan, et al.
Published: (2025)
by: Li, Yinan, et al.
Published: (2025)
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search
by: Nguyen, Tuan Ngo, et al.
Published: (2024)
by: Nguyen, Tuan Ngo, et al.
Published: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
$\varepsilon$-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
by: Li, Yinan, et al.
Published: (2026)
by: Li, Yinan, et al.
Published: (2026)
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
by: Jang, Kyoungseok, et al.
Published: (2024)
by: Jang, Kyoungseok, et al.
Published: (2024)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
by: Qin, Hao, et al.
Published: (2023)
by: Qin, Hao, et al.
Published: (2023)
Mirror Mean-Field Langevin Dynamics
by: Gu, Anming, et al.
Published: (2025)
by: Gu, Anming, et al.
Published: (2025)
$t^3$-Variational Autoencoder: Learning Heavy-tailed Data with Student's t and Power Divergence
by: Kim, Juno, et al.
Published: (2023)
by: Kim, Juno, et al.
Published: (2023)
MEC: Machine-Learning-Assisted Generalized Entropy Calibration for Semi-Supervised Mean Estimation
by: Lee, Se Yoon, et al.
Published: (2026)
by: Lee, Se Yoon, et al.
Published: (2026)
Transformers are Minimax Optimal Nonparametric In-Context Learners
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
by: Qin, Hao, et al.
Published: (2025)
by: Qin, Hao, et al.
Published: (2025)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Learning Explainable Dense Reward Shapes via Bayesian Optimization
by: Koo, Ryan, et al.
Published: (2025)
by: Koo, Ryan, et al.
Published: (2025)
A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets
by: Kubo, Akihiro, et al.
Published: (2026)
by: Kubo, Akihiro, et al.
Published: (2026)
Convergence Rates of Constrained Expected Improvement
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Adaptive Experimentation When You Can't Experiment
by: Zhao, Yao, et al.
Published: (2024)
by: Zhao, Yao, et al.
Published: (2024)
HiPPO-KAN: Efficient KAN Model for Time Series Analysis
by: Lee, SangJong, et al.
Published: (2024)
by: Lee, SangJong, et al.
Published: (2024)
TEDDY: Trimming Edges with Degree-based Discrimination strategY
by: Seo, Hyunjin, et al.
Published: (2024)
by: Seo, Hyunjin, et al.
Published: (2024)
Preference Alignment with Flow Matching
by: Kim, Minu, et al.
Published: (2024)
by: Kim, Minu, et al.
Published: (2024)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
by: Yamamoto, Naoya, et al.
Published: (2025)
by: Yamamoto, Naoya, et al.
Published: (2025)
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
by: Balagopalan, Kapilan, et al.
Published: (2024)
by: Balagopalan, Kapilan, et al.
Published: (2024)
Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
by: Kim, Jeonghye, et al.
Published: (2024)
by: Kim, Jeonghye, et al.
Published: (2024)
Understanding active learning of molecular docking and its applications
by: Kim, Jeonghyeon, et al.
Published: (2024)
by: Kim, Jeonghyeon, et al.
Published: (2024)
Similar Items
-
Beyond RLHF: A Unified Theoretical Framework of Alignment
by: Yun, Jihun, et al.
Published: (2025) -
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
by: Zhao, Yao, et al.
Published: (2026) -
Regularized Online RLHF with Generalized Bilinear Preferences
by: Lee, Junghyun, et al.
Published: (2026) -
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
by: Lee, Junghyun, et al.
Published: (2024) -
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
by: Lee, Junghyun, et al.
Published: (2023)