Oracle-Robust Online Alignment for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zimeng, Gaur, Mudit, Aggarwal, Vaneet |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discrete State Diffusion Models: A Sample Complexity Perspective
von: Srikanth, Aadithya, et al.
Veröffentlicht: (2025)
von: Srikanth, Aadithya, et al.
Veröffentlicht: (2025)
On The Global Convergence Of Online RLHF With Neural Parametrization
von: Gaur, Mudit, et al.
Veröffentlicht: (2024)
von: Gaur, Mudit, et al.
Veröffentlicht: (2024)
Order-Optimal Sample Complexity of Rectified Flows
von: Sahoo, Hari Krishna, et al.
Veröffentlicht: (2026)
von: Sahoo, Hari Krishna, et al.
Veröffentlicht: (2026)
Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
von: Gaur, Mudit, et al.
Veröffentlicht: (2024)
von: Gaur, Mudit, et al.
Veröffentlicht: (2024)
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
Lipschitz Dueling Bandits over Continuous Action Spaces
von: Sharma, Mudit, et al.
Veröffentlicht: (2026)
von: Sharma, Mudit, et al.
Veröffentlicht: (2026)
A Unified Framework for Analyzing Meta-algorithms in Online Convex Optimization
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
Distributionally Robust Self Paced Curriculum Reinforcement Learning
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
Improving Molecule Generation and Drug Discovery with a Knowledge-enhanced Generative Model
von: Malusare, Aditya, et al.
Veröffentlicht: (2024)
von: Malusare, Aditya, et al.
Veröffentlicht: (2024)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning
von: Ganesh, Swetha, et al.
Veröffentlicht: (2026)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2026)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
Efficient $Q$-Learning and Actor-Critic Methods for Robust Average Reward Reinforcement Learning
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Decentralized Projection-free Online Upper-Linearizable Optimization with Applications to DR-Submodular Optimization
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients
von: Xu, Yang, et al.
Veröffentlicht: (2026)
von: Xu, Yang, et al.
Veröffentlicht: (2026)
Sample Complexity Analysis for Constrained Bilevel Reinforcement Learning
von: Saxena, Naman, et al.
Veröffentlicht: (2026)
von: Saxena, Naman, et al.
Veröffentlicht: (2026)
$γ$-weakly $θ$-up-concavity: A Unified Framework for Non-Convex Optimization Beyond DR-Submodular and OSS Functions
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2026)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2026)
From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
Accelerating Quantum Reinforcement Learning with a Quantum Natural Policy Gradient Based Approach
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
Sample-Efficient Constrained Reinforcement Learning with General Parameterization
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2023)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2023)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
Contrastive Cross-Modal Learning for Infusing Chest X-ray Knowledge into ECGs
von: Punyamoorty, Vineet, et al.
Veröffentlicht: (2025)
von: Punyamoorty, Vineet, et al.
Veröffentlicht: (2025)
Hierarchical Deep Counterfactual Regret Minimization
von: Chen, Jiayu, et al.
Veröffentlicht: (2023)
von: Chen, Jiayu, et al.
Veröffentlicht: (2023)
Constrained Reinforcement Learning with Average Reward Objective: Model-Based and Model-Free Algorithms
von: Aggarwal, Vaneet, et al.
Veröffentlicht: (2024)
von: Aggarwal, Vaneet, et al.
Veröffentlicht: (2024)
Stochastic Q-learning for Large Discrete Action Spaces
von: Fourati, Fares, et al.
Veröffentlicht: (2024)
von: Fourati, Fares, et al.
Veröffentlicht: (2024)
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
Upper-Linearizability of Online Non-Monotone DR-Submodular Maximization over Down-Closed Convex Sets
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
Variational Offline Multi-agent Skill Discovery
von: Chen, Jiayu, et al.
Veröffentlicht: (2024)
von: Chen, Jiayu, et al.
Veröffentlicht: (2024)
Understanding the Natural Language of DNA using Encoder-Decoder Foundation Models with Byte-level Precision
von: Malusare, Aditya, et al.
Veröffentlicht: (2023)
von: Malusare, Aditya, et al.
Veröffentlicht: (2023)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
Multi-Agent Combinatorial-Multi-Armed-Bandit framework for the Submodular Welfare Problem under Bandit Feedback
von: Pokhriyal, Subham, et al.
Veröffentlicht: (2026)
von: Pokhriyal, Subham, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Discrete State Diffusion Models: A Sample Complexity Perspective
von: Srikanth, Aadithya, et al.
Veröffentlicht: (2025) -
On The Global Convergence Of Online RLHF With Neural Parametrization
von: Gaur, Mudit, et al.
Veröffentlicht: (2024) -
Order-Optimal Sample Complexity of Rectified Flows
von: Sahoo, Hari Krishna, et al.
Veröffentlicht: (2026) -
Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
von: Gaur, Mudit, et al.
Veröffentlicht: (2024) -
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)