Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Lili, Evans, Ben, Islam, Riashat, Seraj, Raihan, Efroni, Yonathan, Lamb, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PcLast: Discovering Plannable Continuous Latent States
by: Koul, Anurag, et al.
Published: (2023)
by: Koul, Anurag, et al.
Published: (2023)
Gradient Free Deep Reinforcement Learning With TabPFN
by: Schiff, David, et al.
Published: (2025)
by: Schiff, David, et al.
Published: (2025)
Learning Latent Dynamic Robust Representations for World Models
by: Sun, Ruixiang, et al.
Published: (2024)
by: Sun, Ruixiang, et al.
Published: (2024)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2025)
by: Galesloot, Maris F. L., et al.
Published: (2025)
Learning Fused State Representations for Control from Multi-View Observations
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees
by: Azeem, Muqsit, et al.
Published: (2024)
by: Azeem, Muqsit, et al.
Published: (2024)
Contextual bandits with entropy-based human feedback
by: Seraj, Raihan, et al.
Published: (2025)
by: Seraj, Raihan, et al.
Published: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025)
by: Wu, Runzhe, et al.
Published: (2025)
When Does Predictive Inverse Dynamics Outperform Behavior Cloning?
by: Schäfer, Lukas, et al.
Published: (2026)
by: Schäfer, Lukas, et al.
Published: (2026)
Heterogeneous Decentralized Diffusion Models
by: Jiang, Zhiying, et al.
Published: (2026)
by: Jiang, Zhiying, et al.
Published: (2026)
Simple Optimizers for Convex Aligned Multi-Objective Optimization
by: Kretzu, Ben, et al.
Published: (2025)
by: Kretzu, Ben, et al.
Published: (2025)
Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do
by: Wald, Yoav, et al.
Published: (2025)
by: Wald, Yoav, et al.
Published: (2025)
Reinforcement Learning for Sequence Design Leveraging Protein Language Models
by: Subramanian, Jithendaraa, et al.
Published: (2024)
by: Subramanian, Jithendaraa, et al.
Published: (2024)
Multi-Environment POMDPs with Finite-Horizon Objectives
by: Brice, Léonard, et al.
Published: (2026)
by: Brice, Léonard, et al.
Published: (2026)
A Finite-State Controller Based Offline Solver for Deterministic POMDPs
by: Schutz, Alex, et al.
Published: (2025)
by: Schutz, Alex, et al.
Published: (2025)
Pearl: A Production-ready Reinforcement Learning Agent
by: Zhu, Zheqing, et al.
Published: (2023)
by: Zhu, Zheqing, et al.
Published: (2023)
Rethinking Transformers in Solving POMDPs
by: Lu, Chenhao, et al.
Published: (2024)
by: Lu, Chenhao, et al.
Published: (2024)
Generic Multi-modal Representation Learning for Network Traffic Analysis
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
Online Planning in POMDPs with State-Requests
by: Avalos, Raphael, et al.
Published: (2024)
by: Avalos, Raphael, et al.
Published: (2024)
Towards Principled Representation Learning from Videos for Reinforcement Learning
by: Misra, Dipendra, et al.
Published: (2024)
by: Misra, Dipendra, et al.
Published: (2024)
Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
by: Samanta, Ankur, et al.
Published: (2025)
by: Samanta, Ankur, et al.
Published: (2025)
ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs
by: Zhang, Yunuo, et al.
Published: (2025)
by: Zhang, Yunuo, et al.
Published: (2025)
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
by: Motwani, Sumeet Ramesh, et al.
Published: (2025)
by: Motwani, Sumeet Ramesh, et al.
Published: (2025)
Finite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning
by: Hudák, David, et al.
Published: (2026)
by: Hudák, David, et al.
Published: (2026)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
by: Lanier, Michael, et al.
Published: (2024)
by: Lanier, Michael, et al.
Published: (2024)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Scalable Policy-Based RL Algorithms for POMDPs
by: Anjarlekar, Ameya, et al.
Published: (2025)
by: Anjarlekar, Ameya, et al.
Published: (2025)
Towards In-Vehicle Multi-Task Facial Attribute Recognition: Investigating Synthetic Data and Vision Foundation Models
by: Seraj, Esmaeil, et al.
Published: (2024)
by: Seraj, Esmaeil, et al.
Published: (2024)
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
by: Shaw, Seiji, et al.
Published: (2026)
by: Shaw, Seiji, et al.
Published: (2026)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
Value of Information and Reward Specification in Active Inference and POMDPs
by: Wei, Ran
Published: (2024)
by: Wei, Ran
Published: (2024)
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
by: Abdulsamad, Hany, et al.
Published: (2025)
by: Abdulsamad, Hany, et al.
Published: (2025)
How to Explore with Belief: State Entropy Maximization in POMDPs
by: Zamboni, Riccardo, et al.
Published: (2024)
by: Zamboni, Riccardo, et al.
Published: (2024)
Solving Collaborative Dec-POMDPs with Deep Reinforcement Learning Heuristics
by: Soffair, Nitsan
Published: (2022)
by: Soffair, Nitsan
Published: (2022)
Generative Modeling of Class Probability for Multi-Modal Representation Learning
by: Shin, Jungkyoo, et al.
Published: (2025)
by: Shin, Jungkyoo, et al.
Published: (2025)
Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
by: Wang, Yongyi, et al.
Published: (2025)
by: Wang, Yongyi, et al.
Published: (2025)
Learning-based estimation of cattle weight gain and its influencing factors
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
Diffusion Model with Representation Alignment for Protein Inverse Folding
by: Wang, Chenglin, et al.
Published: (2024)
by: Wang, Chenglin, et al.
Published: (2024)
Similar Items
-
PcLast: Discovering Plannable Continuous Latent States
by: Koul, Anurag, et al.
Published: (2023) -
Gradient Free Deep Reinforcement Learning With TabPFN
by: Schiff, David, et al.
Published: (2025) -
Learning Latent Dynamic Robust Representations for World Models
by: Sun, Ruixiang, et al.
Published: (2024) -
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2025) -
Learning Fused State Representations for Control from Multi-View Observations
by: Wang, Zeyu, et al.
Published: (2025)