Gespeichert in:
| Hauptverfasser: | Xiong, Zhihan, Fazel, Maryam, Xiao, Lin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.01249 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits
von: Maynard-Zhang, Leo, et al.
Veröffentlicht: (2026)
von: Maynard-Zhang, Leo, et al.
Veröffentlicht: (2026)
A/B Testing and Best-arm Identification for Linear Bandits with Robustness to Non-stationarity
von: Xiong, Zhihan, et al.
Veröffentlicht: (2023)
von: Xiong, Zhihan, et al.
Veröffentlicht: (2023)
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
von: Bose, Avinandan, et al.
Veröffentlicht: (2025)
von: Bose, Avinandan, et al.
Veröffentlicht: (2025)
Offline congestion games: How feedback type affects data coverage requirement
von: Jiang, Haozhe, et al.
Veröffentlicht: (2022)
von: Jiang, Haozhe, et al.
Veröffentlicht: (2022)
A Black-box Approach for Non-stationary Multi-agent Reinforcement Learning
von: Jiang, Haozhe, et al.
Veröffentlicht: (2023)
von: Jiang, Haozhe, et al.
Veröffentlicht: (2023)
Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
von: Zhou, Runlong, et al.
Veröffentlicht: (2025)
von: Zhou, Runlong, et al.
Veröffentlicht: (2025)
Network-Constrained Policy Optimization for Adaptive Multi-agent Vehicle Routing
von: Arasteh, Fazel, et al.
Veröffentlicht: (2025)
von: Arasteh, Fazel, et al.
Veröffentlicht: (2025)
Toward Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixture Models
von: Xu, Weihang, et al.
Veröffentlicht: (2024)
von: Xu, Weihang, et al.
Veröffentlicht: (2024)
Local linear convergence of gradient methods for overparameterized Gaussian mixtures
von: Wang, Jingxing, et al.
Veröffentlicht: (2026)
von: Wang, Jingxing, et al.
Veröffentlicht: (2026)
Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures
von: Zhou, Mo, et al.
Veröffentlicht: (2025)
von: Zhou, Mo, et al.
Veröffentlicht: (2025)
Offline Multi-task Transfer RL with Representational Penalization
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
von: Bose, Avinandan, et al.
Veröffentlicht: (2025)
von: Bose, Avinandan, et al.
Veröffentlicht: (2025)
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
von: Zhu, Libin, et al.
Veröffentlicht: (2026)
von: Zhu, Libin, et al.
Veröffentlicht: (2026)
Learning Optimal Tax Design in Nonatomic Congestion Games
von: Cui, Qiwen, et al.
Veröffentlicht: (2024)
von: Cui, Qiwen, et al.
Veröffentlicht: (2024)
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
Unregularized Linear Convergence in Zero-Sum Game from Preference Feedback
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
Iterative Linear Quadratic Optimization for Nonlinear Control: Differentiable Programming Algorithmic Templates
von: Roulet, Vincent, et al.
Veröffentlicht: (2022)
von: Roulet, Vincent, et al.
Veröffentlicht: (2022)
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
von: Narang, Adhyyan, et al.
Veröffentlicht: (2026)
von: Narang, Adhyyan, et al.
Veröffentlicht: (2026)
Iteratively reweighted kernel machines efficiently learn sparse functions
von: Zhu, Libin, et al.
Veröffentlicht: (2025)
von: Zhu, Libin, et al.
Veröffentlicht: (2025)
Finite Sample Identification of Partially Observed Bilinear Dynamical Systems
von: Sattar, Yahya, et al.
Veröffentlicht: (2025)
von: Sattar, Yahya, et al.
Veröffentlicht: (2025)
Online SuBmodular + SuPermodular (BP) Maximization with Bandit Feedback
von: Narang, Adhyyan, et al.
Veröffentlicht: (2022)
von: Narang, Adhyyan, et al.
Veröffentlicht: (2022)
Convergence Dynamics of Over-Parameterized Score Matching for a Single Gaussian
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
High-dimensional Limit of SGD for Diagonal Linear Networks
von: Malaxechebarría, Begoña García, et al.
Veröffentlicht: (2026)
von: Malaxechebarría, Begoña García, et al.
Veröffentlicht: (2026)
Optimization and generalization analysis for two-layer physics-informed neural networks without over-parametrization
von: Zeng, Zhihan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhihan, et al.
Veröffentlicht: (2025)
Global Convergence of Four-Layer Matrix Factorization under Random Initialization
von: Luo, Minrui, et al.
Veröffentlicht: (2025)
von: Luo, Minrui, et al.
Veröffentlicht: (2025)
Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics
von: Choi, Sunmook, et al.
Veröffentlicht: (2025)
von: Choi, Sunmook, et al.
Veröffentlicht: (2025)
Sub-optimality of the Separation Principle for Quadratic Control from Bilinear Observations
von: Sattar, Yahya, et al.
Veröffentlicht: (2025)
von: Sattar, Yahya, et al.
Veröffentlicht: (2025)
AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
Emergent specialization from participation dynamics and multi-learner retraining
von: Dean, Sarah, et al.
Veröffentlicht: (2022)
von: Dean, Sarah, et al.
Veröffentlicht: (2022)
Improving Credit Card Fraud Detection with an Optimized Explainable Boosting Machine
von: Fazel, Reza E., et al.
Veröffentlicht: (2026)
von: Fazel, Reza E., et al.
Veröffentlicht: (2026)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
Divergence-Augmented Policy Optimization
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Federated Offline Policy Optimization with Dual Regularization
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
Primal-Dual Policy Optimization for Linear CMDPs with Adversarial Losses
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
Universal Approximation of Operators with Transformers and Neural Integral Operators
von: Zappala, Emanuele, et al.
Veröffentlicht: (2024)
von: Zappala, Emanuele, et al.
Veröffentlicht: (2024)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization Lens
von: Qiu, Junbin, et al.
Veröffentlicht: (2026)
von: Qiu, Junbin, et al.
Veröffentlicht: (2026)
Soft Adaptive Policy Optimization
von: Gao, Chang, et al.
Veröffentlicht: (2025)
von: Gao, Chang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
von: Bose, Avinandan, et al.
Veröffentlicht: (2024) -
On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits
von: Maynard-Zhang, Leo, et al.
Veröffentlicht: (2026) -
A/B Testing and Best-arm Identification for Linear Bandits with Robustness to Non-stationarity
von: Xiong, Zhihan, et al.
Veröffentlicht: (2023) -
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
von: Bose, Avinandan, et al.
Veröffentlicht: (2025) -
Offline congestion games: How feedback type affects data coverage requirement
von: Jiang, Haozhe, et al.
Veröffentlicht: (2022)