FraPPE: Fast and Efficient Preference-based Pure Exploration
Fuente:
arXiv
Guardado en:
| Autores principales: | Das, Udvas, Shukla, Apurv, Basu, Debabrota |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
por: Basu, Debabrota, et al.
Publicado: (2025)
por: Basu, Debabrota, et al.
Publicado: (2025)
Preference-based Pure Exploration
por: Shukla, Apurv, et al.
Publicado: (2024)
por: Shukla, Apurv, et al.
Publicado: (2024)
Learning to Explore with Lagrangians for Bandits under Unknown Linear Constraints
por: Das, Udvas, et al.
Publicado: (2024)
por: Das, Udvas, et al.
Publicado: (2024)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
por: Mohamed, Mimoun, et al.
Publicado: (2023)
por: Mohamed, Mimoun, et al.
Publicado: (2023)
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
por: Kitaoka, Akira
Publicado: (2025)
por: Kitaoka, Akira
Publicado: (2025)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
por: Han, Qiyang, et al.
Publicado: (2025)
por: Han, Qiyang, et al.
Publicado: (2025)
Learning to Fuse Temporal Proximity Networks: A Case Study in Chimpanzee Social Interactions
por: He, Yixuan, et al.
Publicado: (2025)
por: He, Yixuan, et al.
Publicado: (2025)
Statistical and Algorithmic Foundations of Reinforcement Learning
por: Chi, Yuejie, et al.
Publicado: (2025)
por: Chi, Yuejie, et al.
Publicado: (2025)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
por: Nguyen, Minh
Publicado: (2026)
por: Nguyen, Minh
Publicado: (2026)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
por: Bareilles, Gilles, et al.
Publicado: (2026)
por: Bareilles, Gilles, et al.
Publicado: (2026)
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
por: Mustafi, Aratrika, et al.
Publicado: (2026)
por: Mustafi, Aratrika, et al.
Publicado: (2026)
A Differential and Pointwise Control Approach to Reinforcement Learning
por: Nguyen, Minh, et al.
Publicado: (2024)
por: Nguyen, Minh, et al.
Publicado: (2024)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
por: Chen, Siyu, et al.
Publicado: (2024)
por: Chen, Siyu, et al.
Publicado: (2024)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
por: Rashidinejad, Paria, et al.
Publicado: (2024)
por: Rashidinejad, Paria, et al.
Publicado: (2024)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
por: Yan, Shunxing, et al.
Publicado: (2026)
por: Yan, Shunxing, et al.
Publicado: (2026)
Smooth Non-Stationary Bandits
por: Jia, Su, et al.
Publicado: (2023)
por: Jia, Su, et al.
Publicado: (2023)
Piecewise Polynomial Regression of Tame Functions via Integer Programming
por: Bareilles, Gilles, et al.
Publicado: (2023)
por: Bareilles, Gilles, et al.
Publicado: (2023)
The Fair Game: Auditing & Debiasing AI Algorithms Over Time
por: Basu, Debabrota, et al.
Publicado: (2025)
por: Basu, Debabrota, et al.
Publicado: (2025)
Differentially Private High Dimensional Bandits
por: Shukla, Apurv
Publicado: (2024)
por: Shukla, Apurv
Publicado: (2024)
Geometry-induced Regularization in Deep ReLU Neural Networks
por: Bona-Pellissier, Joachim, et al.
Publicado: (2024)
por: Bona-Pellissier, Joachim, et al.
Publicado: (2024)
Efficient Group Lasso Regularized Rank Regression with Data-Driven Parameter Determination
por: Lin, Meixia, et al.
Publicado: (2025)
por: Lin, Meixia, et al.
Publicado: (2025)
Data-Efficient Non-Gaussian Semi-Nonparametric Density Estimation for Nonlinear Dynamical Systems
por: Liao, Aaron R., et al.
Publicado: (2026)
por: Liao, Aaron R., et al.
Publicado: (2026)
Fast convergence of the Expectation Maximization algorithm under a logarithmic Sobolev inequality
por: Caprio, Rocco, et al.
Publicado: (2024)
por: Caprio, Rocco, et al.
Publicado: (2024)
Convergence of flow-based generative models via proximal gradient descent in Wasserstein space
por: Cheng, Xiuyuan, et al.
Publicado: (2023)
por: Cheng, Xiuyuan, et al.
Publicado: (2023)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
por: Thekumparampil, Kiran Koshy, et al.
Publicado: (2024)
por: Thekumparampil, Kiran Koshy, et al.
Publicado: (2024)
Fast Spawn\&Prune (FS\&P): Global convergence of stochastic conic particle gradient descent via birth/death process
por: De Castro, Yohann, et al.
Publicado: (2026)
por: De Castro, Yohann, et al.
Publicado: (2026)
Analytic Bridge Diffusions for Controlled Path Generation
por: Chertkov, Michael
Publicado: (2026)
por: Chertkov, Michael
Publicado: (2026)
Causal Invariance Learning via Efficient Nonconvex Optimization
por: Wang, Zhenyu, et al.
Publicado: (2024)
por: Wang, Zhenyu, et al.
Publicado: (2024)
A Graphical Global Optimization Framework for Parameter Estimation of Statistical Models with Nonconvex Regularization Functions
por: Davarnia, Danial, et al.
Publicado: (2025)
por: Davarnia, Danial, et al.
Publicado: (2025)
ECPv2: Fast, Efficient, and Scalable Global Optimization of Lipschitz Functions
por: Fourati, Fares, et al.
Publicado: (2025)
por: Fourati, Fares, et al.
Publicado: (2025)
Delightful Exploration
por: Osband, Ian
Publicado: (2026)
por: Osband, Ian
Publicado: (2026)
Stochastic Optimization with Optimal Importance Sampling
por: Aolaritei, Liviu, et al.
Publicado: (2025)
por: Aolaritei, Liviu, et al.
Publicado: (2025)
Joint learning of a network of linear dynamical systems via total variation penalization
por: Donnat, Claire, et al.
Publicado: (2025)
por: Donnat, Claire, et al.
Publicado: (2025)
A review of NMF, PLSA, LBA, EMA, and LCA with a focus on the identifiability issue
por: Qi, Qianqian, et al.
Publicado: (2025)
por: Qi, Qianqian, et al.
Publicado: (2025)
Error Analysis of Triangular Optimal Transport Maps for Filtering
por: Al-Jarrah, Mohammad, et al.
Publicado: (2025)
por: Al-Jarrah, Mohammad, et al.
Publicado: (2025)
Online Inference of Constrained Optimization: Primal-Dual Optimality and Sequential Quadratic Programming
por: Gao, Yihang, et al.
Publicado: (2025)
por: Gao, Yihang, et al.
Publicado: (2025)
Mixing Times and Privacy Analysis for the Projected Langevin Algorithm under a Modulus of Continuity
por: Bravo, Mario, et al.
Publicado: (2025)
por: Bravo, Mario, et al.
Publicado: (2025)
Extreme mass distributions for quasi-copulas
por: Omladič, Matjaž, et al.
Publicado: (2025)
por: Omladič, Matjaž, et al.
Publicado: (2025)
Learning an Optimal Assortment Policy under Observational Data
por: Han, Yuxuan, et al.
Publicado: (2025)
por: Han, Yuxuan, et al.
Publicado: (2025)
An Elementary Proof of the Near Optimality of LogSumExp Smoothing
por: Samakhoana, Thabo, et al.
Publicado: (2025)
por: Samakhoana, Thabo, et al.
Publicado: (2025)
Ejemplares similares
-
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
por: Basu, Debabrota, et al.
Publicado: (2025) -
Preference-based Pure Exploration
por: Shukla, Apurv, et al.
Publicado: (2024) -
Learning to Explore with Lagrangians for Bandits under Unknown Linear Constraints
por: Das, Udvas, et al.
Publicado: (2024) -
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
por: Mohamed, Mimoun, et al.
Publicado: (2023) -
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
por: Kitaoka, Akira
Publicado: (2025)