Provably Efficient Interactive-Grounded Learning with Personalized Reward
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Mengxiao, Zhang, Yuheng, Luo, Haipeng, Mineiro, Paul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
by: Zhang, Mengxiao, et al.
Published: (2026)
by: Zhang, Mengxiao, et al.
Published: (2026)
Efficient Contextual Bandits with Uninformed Feedback Graphs
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Contextual Multinomial Logit Bandits with General Value Functions
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Contextual Linear Bandits with Delay as Payoff
by: Zhang, Mengxiao, et al.
Published: (2025)
by: Zhang, Mengxiao, et al.
Published: (2025)
Alternating Regret for Online Convex Optimization
by: Hait, Soumita, et al.
Published: (2025)
by: Hait, Soumita, et al.
Published: (2025)
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
by: Hait, Soumita, et al.
Published: (2026)
by: Hait, Soumita, et al.
Published: (2026)
Comparator-Adaptive $Φ$-Regret: Improved Bounds, Simpler Algorithms, and Applications to Games
by: Hait, Soumita, et al.
Published: (2025)
by: Hait, Soumita, et al.
Published: (2025)
Provably Sample-Efficient Robust Reinforcement Learning with Average Reward
by: Roch, Zachary, et al.
Published: (2025)
by: Roch, Zachary, et al.
Published: (2025)
No-Regret Learning for Fair Multi-Agent Social Welfare Optimization
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Online Joint Fine-tuning of Multi-Agent Flows
by: Mineiro, Paul
Published: (2024)
by: Mineiro, Paul
Published: (2024)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
by: Vora, Kevin, et al.
Published: (2025)
by: Vora, Kevin, et al.
Published: (2025)
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning
by: Deng, Yihe, et al.
Published: (2024)
by: Deng, Yihe, et al.
Published: (2024)
Provably Efficient Online RLHF with One-Pass Reward Modeling
by: Li, Long-Fei, et al.
Published: (2025)
by: Li, Long-Fei, et al.
Published: (2025)
Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Provably Efficient Exploration in Reward Machines with Low Regret
by: Bourel, Hippolyte, et al.
Published: (2024)
by: Bourel, Hippolyte, et al.
Published: (2024)
Active, anytime-valid risk controlling prediction sets
by: Xu, Ziyu, et al.
Published: (2024)
by: Xu, Ziyu, et al.
Published: (2024)
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
by: Feng, Youhe, et al.
Published: (2026)
by: Feng, Youhe, et al.
Published: (2026)
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback
by: Pang, Lei, et al.
Published: (2025)
by: Pang, Lei, et al.
Published: (2025)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
A Short Note on a Variant of the Squint Algorithm
by: Luo, Haipeng
Published: (2026)
by: Luo, Haipeng
Published: (2026)
Provably Efficient Partially Observable Risk-Sensitive Reinforcement Learning with Hindsight Observation
by: Zhang, Tonghe, et al.
Published: (2024)
by: Zhang, Tonghe, et al.
Published: (2024)
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
by: Xu, Tian, et al.
Published: (2024)
by: Xu, Tian, et al.
Published: (2024)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Provable Sample-Efficient Transfer Learning Conditional Diffusion Models via Representation Learning
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
On Provable Benefits of Muon in Federated Learning
by: Zhang, Xinwen, et al.
Published: (2025)
by: Zhang, Xinwen, et al.
Published: (2025)
Provable Multi-Task Reinforcement Learning: A Representation Learning Framework with Low Rank Rewards
by: Guo, Yaoze, et al.
Published: (2026)
by: Guo, Yaoze, et al.
Published: (2026)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
by: Ma, Qiyao, et al.
Published: (2026)
by: Ma, Qiyao, et al.
Published: (2026)
Provably Efficient and Agile Randomized Q-Learning
by: Wang, He, et al.
Published: (2025)
by: Wang, He, et al.
Published: (2025)
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning
by: Luo, Zhi, et al.
Published: (2024)
by: Luo, Zhi, et al.
Published: (2024)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2024)
by: Qiu, Shuang, et al.
Published: (2024)
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
by: Zhang, Yuheng, et al.
Published: (2026)
by: Zhang, Yuheng, et al.
Published: (2026)
Efficient Reinforcement Learning in Probabilistic Reward Machines
by: Lin, Xiaofeng, et al.
Published: (2024)
by: Lin, Xiaofeng, et al.
Published: (2024)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
by: Bu, Dake, et al.
Published: (2024)
by: Bu, Dake, et al.
Published: (2024)
Exploiting Curvature in Online Convex Optimization with Delayed Feedback
by: Qiu, Hao, et al.
Published: (2025)
by: Qiu, Hao, et al.
Published: (2025)
An Efficient Subgraph GNN with Provable Substructure Counting Power
by: Yan, Zuoyu, et al.
Published: (2023)
by: Yan, Zuoyu, et al.
Published: (2023)
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
Eidetic Learning: an Efficient and Provable Solution to Catastrophic Forgetting
by: Dronen, Nicholas, et al.
Published: (2025)
by: Dronen, Nicholas, et al.
Published: (2025)
Similar Items
-
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
by: Zhang, Mengxiao, et al.
Published: (2026) -
Efficient Contextual Bandits with Uninformed Feedback Graphs
by: Zhang, Mengxiao, et al.
Published: (2024) -
Contextual Multinomial Logit Bandits with General Value Functions
by: Zhang, Mengxiao, et al.
Published: (2024) -
Contextual Linear Bandits with Delay as Payoff
by: Zhang, Mengxiao, et al.
Published: (2025) -
Alternating Regret for Online Convex Optimization
by: Hait, Soumita, et al.
Published: (2025)