Value of Information and Reward Specification in Active Inference and POMDPs
Fuente:
arXiv
Salvato in:
| Autore principale: | Wei, Ran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Transformers in Solving POMDPs
di: Lu, Chenhao, et al.
Pubblicazione: (2024)
di: Lu, Chenhao, et al.
Pubblicazione: (2024)
Active Inference with Reusable State-Dependent Value Profiles
di: Poschl, Jacob
Pubblicazione: (2025)
di: Poschl, Jacob
Pubblicazione: (2025)
Online Planning in POMDPs with State-Requests
di: Avalos, Raphael, et al.
Pubblicazione: (2024)
di: Avalos, Raphael, et al.
Pubblicazione: (2024)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
di: Galesloot, Maris F. L., et al.
Pubblicazione: (2024)
di: Galesloot, Maris F. L., et al.
Pubblicazione: (2024)
Scalable Policy-Based RL Algorithms for POMDPs
di: Anjarlekar, Ameya, et al.
Pubblicazione: (2025)
di: Anjarlekar, Ameya, et al.
Pubblicazione: (2025)
Learning Logic Specifications for Policy Guidance in POMDPs: an Inductive Logic Programming Approach
di: Meli, Daniele, et al.
Pubblicazione: (2024)
di: Meli, Daniele, et al.
Pubblicazione: (2024)
How to Explore with Belief: State Entropy Maximization in POMDPs
di: Zamboni, Riccardo, et al.
Pubblicazione: (2024)
di: Zamboni, Riccardo, et al.
Pubblicazione: (2024)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
di: Wendland, Joshua, et al.
Pubblicazione: (2026)
di: Wendland, Joshua, et al.
Pubblicazione: (2026)
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
di: Abdulsamad, Hany, et al.
Pubblicazione: (2025)
di: Abdulsamad, Hany, et al.
Pubblicazione: (2025)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
di: Huang, Chenghua, et al.
Pubblicazione: (2025)
di: Huang, Chenghua, et al.
Pubblicazione: (2025)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
Tiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
di: Zhou, Zhiyuan, et al.
Pubblicazione: (2022)
di: Zhou, Zhiyuan, et al.
Pubblicazione: (2022)
Transductive Reward Inference on Graph
di: Qu, Bohao, et al.
Pubblicazione: (2024)
di: Qu, Bohao, et al.
Pubblicazione: (2024)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
di: Galesloot, Maris F. L., et al.
Pubblicazione: (2025)
di: Galesloot, Maris F. L., et al.
Pubblicazione: (2025)
Quasimetric Value Functions with Dense Rewards
di: Valieva, Khadichabonu, et al.
Pubblicazione: (2024)
di: Valieva, Khadichabonu, et al.
Pubblicazione: (2024)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
di: Zhang, Yuheng, et al.
Pubblicazione: (2025)
di: Zhang, Yuheng, et al.
Pubblicazione: (2025)
Learning An Active Inference Model of Driver Perception and Control: Application to Vehicle Car-Following
di: Wei, Ran, et al.
Pubblicazione: (2023)
di: Wei, Ran, et al.
Pubblicazione: (2023)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
di: Wu, Lili, et al.
Pubblicazione: (2024)
di: Wu, Lili, et al.
Pubblicazione: (2024)
Value-Free Policy Optimization via Reward Partitioning
di: Faye, Bilal, et al.
Pubblicazione: (2025)
di: Faye, Bilal, et al.
Pubblicazione: (2025)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
di: Lanier, Michael, et al.
Pubblicazione: (2024)
di: Lanier, Michael, et al.
Pubblicazione: (2024)
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
di: Shaw, Seiji, et al.
Pubblicazione: (2026)
di: Shaw, Seiji, et al.
Pubblicazione: (2026)
Efficient Process Reward Model Training via Active Learning
di: Duan, Keyu, et al.
Pubblicazione: (2025)
di: Duan, Keyu, et al.
Pubblicazione: (2025)
Solving Collaborative Dec-POMDPs with Deep Reinforcement Learning Heuristics
di: Soffair, Nitsan
Pubblicazione: (2022)
di: Soffair, Nitsan
Pubblicazione: (2022)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
di: Wang, Haichuan, et al.
Pubblicazione: (2026)
di: Wang, Haichuan, et al.
Pubblicazione: (2026)
Active Timepoint Selection for Learning Measure-Valued Trajectories
di: Huynh, Nicolas, et al.
Pubblicazione: (2026)
di: Huynh, Nicolas, et al.
Pubblicazione: (2026)
ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs
di: Zhang, Yunuo, et al.
Pubblicazione: (2025)
di: Zhang, Yunuo, et al.
Pubblicazione: (2025)
PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models
di: Deng, Fei, et al.
Pubblicazione: (2024)
di: Deng, Fei, et al.
Pubblicazione: (2024)
Expressive Temporal Specifications for Reward Monitoring
di: Adalat, Omar, et al.
Pubblicazione: (2025)
di: Adalat, Omar, et al.
Pubblicazione: (2025)
Batch Active Learning of Reward Functions from Human Preferences
di: Bıyık, Erdem, et al.
Pubblicazione: (2024)
di: Bıyık, Erdem, et al.
Pubblicazione: (2024)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
Shapley-Value-Based Graph Sparsification for GNN Inference
di: Akkas, Selahattin, et al.
Pubblicazione: (2025)
di: Akkas, Selahattin, et al.
Pubblicazione: (2025)
Inference-Time Scaling for Generalist Reward Modeling
di: Liu, Zijun, et al.
Pubblicazione: (2025)
di: Liu, Zijun, et al.
Pubblicazione: (2025)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
di: Baur, Raphaël, et al.
Pubblicazione: (2026)
di: Baur, Raphaël, et al.
Pubblicazione: (2026)
Value Internalization: Learning and Generalizing from Social Reward
di: Rong, Frieda, et al.
Pubblicazione: (2024)
di: Rong, Frieda, et al.
Pubblicazione: (2024)
Deep Active Inference Agents for Delayed and Long-Horizon Environments
di: Yeganeh, Yavar Taheri, et al.
Pubblicazione: (2025)
di: Yeganeh, Yavar Taheri, et al.
Pubblicazione: (2025)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
di: Zhu, Xiao, et al.
Pubblicazione: (2026)
di: Zhu, Xiao, et al.
Pubblicazione: (2026)
Solving Truly Massive Budgeted Monotonic POMDPs with Oracle-Guided Meta-Reinforcement Learning
di: Vora, Manav, et al.
Pubblicazione: (2024)
di: Vora, Manav, et al.
Pubblicazione: (2024)
Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Rethinking Transformers in Solving POMDPs
di: Lu, Chenhao, et al.
Pubblicazione: (2024) -
Active Inference with Reusable State-Dependent Value Profiles
di: Poschl, Jacob
Pubblicazione: (2025) -
Online Planning in POMDPs with State-Requests
di: Avalos, Raphael, et al.
Pubblicazione: (2024) -
Pessimistic Iterative Planning with RNNs for Robust POMDPs
di: Galesloot, Maris F. L., et al.
Pubblicazione: (2024) -
Scalable Policy-Based RL Algorithms for POMDPs
di: Anjarlekar, Ameya, et al.
Pubblicazione: (2025)