Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chaudhari, Shreyas, Deshpande, Ameet, da Silva, Bruno Castro, Thomas, Philip S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2025)
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2025)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
PersonaGym: Evaluating Persona Agents and LLMs
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
Agent Context Protocols Enhance Collective Inference
von: Bhardwaj, Devansh, et al.
Veröffentlicht: (2025)
von: Bhardwaj, Devansh, et al.
Veröffentlicht: (2025)
Low Variance Off-policy Evaluation with State-based Importance Sampling
von: Bossens, David M., et al.
Veröffentlicht: (2022)
von: Bossens, David M., et al.
Veröffentlicht: (2022)
Policy Gradient Methods in the Presence of Symmetries and State Abstractions
von: Panangaden, Prakash, et al.
Veröffentlicht: (2023)
von: Panangaden, Prakash, et al.
Veröffentlicht: (2023)
A Unifying View of Coverage in Linear Off-Policy Evaluation
von: Amortila, Philip, et al.
Veröffentlicht: (2026)
von: Amortila, Philip, et al.
Veröffentlicht: (2026)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
von: Bennett, Andrew, et al.
Veröffentlicht: (2024)
von: Bennett, Andrew, et al.
Veröffentlicht: (2024)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
von: Sun, Hao, et al.
Veröffentlicht: (2023)
von: Sun, Hao, et al.
Veröffentlicht: (2023)
Conformal Off-Policy Evaluation in Markov Decision Processes
von: Foffano, Daniele, et al.
Veröffentlicht: (2023)
von: Foffano, Daniele, et al.
Veröffentlicht: (2023)
Concept-driven Off Policy Evaluation
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
Clustering Context in Off-Policy Evaluation
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
Learning Action Embeddings for Off-Policy Evaluation
von: Cief, Matej, et al.
Veröffentlicht: (2023)
von: Cief, Matej, et al.
Veröffentlicht: (2023)
Tiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2022)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2022)
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
On the Sample Efficiency of Abstractions and Potential-Based Reward Shaping in Reinforcement Learning
von: Canonaco, Giuseppe, et al.
Veröffentlicht: (2024)
von: Canonaco, Giuseppe, et al.
Veröffentlicht: (2024)
Learning Consistent Causal Abstraction Networks
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2026)
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
Learning with Language-Guided State Abstractions
von: Peng, Andi, et al.
Veröffentlicht: (2024)
von: Peng, Andi, et al.
Veröffentlicht: (2024)
Leveraging Machine Learning for Early Autism Detection via INDT-ASD Indian Database
von: Shrivastava, Trapti, et al.
Veröffentlicht: (2024)
von: Shrivastava, Trapti, et al.
Veröffentlicht: (2024)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
von: Shimizu, Tatsuhiro, et al.
Veröffentlicht: (2024)
von: Shimizu, Tatsuhiro, et al.
Veröffentlicht: (2024)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
von: Aouali, Imad, et al.
Veröffentlicht: (2024)
von: Aouali, Imad, et al.
Veröffentlicht: (2024)
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
von: Shimizu, Tatsuhiro, et al.
Veröffentlicht: (2025)
von: Shimizu, Tatsuhiro, et al.
Veröffentlicht: (2025)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement Learning
von: Azran, Guy, et al.
Veröffentlicht: (2023)
von: Azran, Guy, et al.
Veröffentlicht: (2023)
Balancing Immediate Revenue and Future Off-Policy Evaluation in Coupon Allocation
von: Nishimura, Naoki, et al.
Veröffentlicht: (2024)
von: Nishimura, Naoki, et al.
Veröffentlicht: (2024)
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2023)
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2023)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
Intra-Trajectory Consistency for Reward Modeling
von: Zhou, Chaoyang, et al.
Veröffentlicht: (2025)
von: Zhou, Chaoyang, et al.
Veröffentlicht: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
von: Fan, Jiajun, et al.
Veröffentlicht: (2025)
von: Fan, Jiajun, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Exogenous States and Rewards
von: Trimponias, George, et al.
Veröffentlicht: (2023)
von: Trimponias, George, et al.
Veröffentlicht: (2023)
From Noise to Control: Parameterized Diffusion Policies
von: Zhang, Renhao, et al.
Veröffentlicht: (2026)
von: Zhang, Renhao, et al.
Veröffentlicht: (2026)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2025) -
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024) -
PersonaGym: Evaluating Persona Agents and LLMs
von: Samuel, Vinay, et al.
Veröffentlicht: (2024) -
Agent Context Protocols Enhance Collective Inference
von: Bhardwaj, Devansh, et al.
Veröffentlicht: (2025) -
Low Variance Off-policy Evaluation with State-based Importance Sampling
von: Bossens, David M., et al.
Veröffentlicht: (2022)