Inferring Reward Machines and Transition Machines from Partially Observable Markov Decision Processes
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yuly, Liu, Jiamou, Zhang, Libo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
Networked Agents in the Dark: Team Value Learning under Partial Observability
by: Varela, Guilherme S., et al.
Published: (2025)
by: Varela, Guilherme S., et al.
Published: (2025)
When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability
by: Haklidir, Mehmet
Published: (2026)
by: Haklidir, Mehmet
Published: (2026)
Manipulating Predictions over Discrete Inputs in Machine Teaching
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
Tailoring Machine Learning for Process Mining
by: Ceravolo, Paolo, et al.
Published: (2023)
by: Ceravolo, Paolo, et al.
Published: (2023)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
OFMU: Optimization-Driven Framework for Machine Unlearning
by: Asif, Sadia, et al.
Published: (2025)
by: Asif, Sadia, et al.
Published: (2025)
Reward Machines for Deep RL in Noisy and Uncertain Environments
by: Li, Andrew C., et al.
Published: (2024)
by: Li, Andrew C., et al.
Published: (2024)
Evaluating Explanatory Capabilities of Machine Learning Models in Medical Diagnostics: A Human-in-the-Loop Approach
by: Bobes-Bascarán, José, et al.
Published: (2024)
by: Bobes-Bascarán, José, et al.
Published: (2024)
Machine Learning vs Deep Learning: The Generalization Problem
by: Bay, Yong Yi, et al.
Published: (2024)
by: Bay, Yong Yi, et al.
Published: (2024)
Extreme Learning Machines for Fast Training of Click-Through Rate Prediction Models
by: Biçici, Ergun
Published: (2024)
by: Biçici, Ergun
Published: (2024)
Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
by: Groeneveld, Jan Niklas, et al.
Published: (2025)
by: Groeneveld, Jan Niklas, et al.
Published: (2025)
Observation Adaptation via Annealed Importance Resampling for Partially Observable Markov Decision Processes
by: Zhang, Yunuo, et al.
Published: (2025)
by: Zhang, Yunuo, et al.
Published: (2025)
Not All Transitions Matter: Evidence from PPO
by: Basnet, Ajhesh
Published: (2026)
by: Basnet, Ajhesh
Published: (2026)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
LakeMLB: Data Lake Machine Learning Benchmark
by: Pan, Feiyu, et al.
Published: (2026)
by: Pan, Feiyu, et al.
Published: (2026)
CroSel: Cross Selection of Confident Pseudo Labels for Partial-Label Learning
by: Tian, Shiyu, et al.
Published: (2023)
by: Tian, Shiyu, et al.
Published: (2023)
Opponent State Inference Under Partial Observability: An HMM-POMDP Framework for 2026 Formula 1 Energy Strategy
by: Kleisarchaki, Kalliopi
Published: (2026)
by: Kleisarchaki, Kalliopi
Published: (2026)
A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences
by: Siow, Bryan Y.
Published: (2025)
by: Siow, Bryan Y.
Published: (2025)
Towards a Robust Soft Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms
by: Alhakami, Mohannad, et al.
Published: (2024)
by: Alhakami, Mohannad, et al.
Published: (2024)
SourceSplice: Source Selection for Machine Learning Tasks
by: Singh, Ambarish, et al.
Published: (2025)
by: Singh, Ambarish, et al.
Published: (2025)
Counterfactual Strategies for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2025)
by: Kobialka, Paul, et al.
Published: (2025)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
Improved Performances and Motivation in Intelligent Tutoring Systems: Combining Machine Learning and Learner Choice
by: Clément, Benjamin, et al.
Published: (2024)
by: Clément, Benjamin, et al.
Published: (2024)
Evaluating the Energy Consumption of Machine Learning: Systematic Literature Review and Experiments
by: Rodriguez, Charlotte, et al.
Published: (2024)
by: Rodriguez, Charlotte, et al.
Published: (2024)
Evaluating SAP RPT-1 for Enterprise Business Process Prediction: In-Context Learning vs. Traditional Machine Learning on Structured SAP Data
by: Lal, Amit
Published: (2026)
by: Lal, Amit
Published: (2026)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
by: Lan, Guangchen, et al.
Published: (2026)
by: Lan, Guangchen, et al.
Published: (2026)
From Numbers to Prompts: A Cognitive Symbolic Transition Mechanism for Lightweight Time-Series Forecasting
by: Yoon, Namkyung, et al.
Published: (2026)
by: Yoon, Namkyung, et al.
Published: (2026)
Attribution-based Explanations for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2026)
by: Kobialka, Paul, et al.
Published: (2026)
Interpreting Machine Learning Models for Room Temperature Prediction in Non-domestic Buildings
by: Mao, Jianqiao, et al.
Published: (2021)
by: Mao, Jianqiao, et al.
Published: (2021)
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
by: Zhang, Junkai, et al.
Published: (2025)
by: Zhang, Junkai, et al.
Published: (2025)
Machine Learning for Energy-Performance-aware Scheduling
by: Hu, Zheyuan, et al.
Published: (2026)
by: Hu, Zheyuan, et al.
Published: (2026)
Machines of Meaning
by: Nunes, Davide, et al.
Published: (2024)
by: Nunes, Davide, et al.
Published: (2024)
Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making
by: Liu, Larkin, et al.
Published: (2025)
by: Liu, Larkin, et al.
Published: (2025)
Machine Learning with a Reject Option: A survey
by: Hendrickx, Kilian, et al.
Published: (2021)
by: Hendrickx, Kilian, et al.
Published: (2021)
Towards Independence Criterion in Machine Unlearning of Features and Labels
by: Han, Ling, et al.
Published: (2024)
by: Han, Ling, et al.
Published: (2024)
What is Reproducibility in Artificial Intelligence and Machine Learning Research?
by: Desai, Abhyuday, et al.
Published: (2024)
by: Desai, Abhyuday, et al.
Published: (2024)
Risk-Sensitive Multi-Agent Reinforcement Learning in Network Aggregative Markov Games
by: Ghaemi, Hafez, et al.
Published: (2024)
by: Ghaemi, Hafez, et al.
Published: (2024)
Interactive Diabetes Risk Prediction Using Explainable Machine Learning: A Dash-Based Approach with SHAP, LIME, and Comorbidity Insights
by: Allani, Udaya
Published: (2025)
by: Allani, Udaya
Published: (2025)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
Similar Items
-
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026) -
Networked Agents in the Dark: Team Value Learning under Partial Observability
by: Varela, Guilherme S., et al.
Published: (2025) -
When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability
by: Haklidir, Mehmet
Published: (2026) -
Manipulating Predictions over Discrete Inputs in Machine Teaching
by: Wu, Xiaodong, et al.
Published: (2024) -
Tailoring Machine Learning for Process Mining
by: Ceravolo, Paolo, et al.
Published: (2023)