The Value of Mechanistic Priors in Sequential Decision Making
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shufaro, Itai, Benor, Gal, Mannor, Shie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Bits and Bandits: Quantifying the Regret-Information Trade-off
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
von: Valensi, David, et al.
Veröffentlicht: (2024)
von: Valensi, David, et al.
Veröffentlicht: (2024)
Policy Gradient with Tree Expansion
von: Dalal, Gal, et al.
Veröffentlicht: (2023)
von: Dalal, Gal, et al.
Veröffentlicht: (2023)
MinMaxMin $Q$-learning
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Conservative DDPG -- Pessimistic RL without Ensemble
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Representative Action Selection for Large Action Space: From Bandits to MDPs
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Segment Feedback
von: Du, Yihan, et al.
Veröffentlicht: (2025)
von: Du, Yihan, et al.
Veröffentlicht: (2025)
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
von: Perets, Binyamin, et al.
Veröffentlicht: (2026)
von: Perets, Binyamin, et al.
Veröffentlicht: (2026)
Efficient Fairness-Performance Pareto Front Computation
von: Kozdoba, Mark, et al.
Veröffentlicht: (2024)
von: Kozdoba, Mark, et al.
Veröffentlicht: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
von: Du, Yihan, et al.
Veröffentlicht: (2024)
von: Du, Yihan, et al.
Veröffentlicht: (2024)
Representation-Driven Reinforcement Learning
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
Sobolev Space Regularised Pre Density Models
von: Kozdoba, Mark, et al.
Veröffentlicht: (2023)
von: Kozdoba, Mark, et al.
Veröffentlicht: (2023)
Representative Action Selection for Large Action Space Bandit Families
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
von: Gadot, Uri, et al.
Veröffentlicht: (2025)
von: Gadot, Uri, et al.
Veröffentlicht: (2025)
Spectral Bellman Method: Unifying Representation and Exploration in RL
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
A Classification View on Meta Learning Bandits
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
Task Tokens: A Flexible Approach to Adapting Behavior Foundation Models
von: Vainshtein, Ron, et al.
Veröffentlicht: (2025)
von: Vainshtein, Ron, et al.
Veröffentlicht: (2025)
Implementing Reinforcement Learning Datacenter Congestion Control in NVIDIA NICs
von: Fuhrer, Benjamin, et al.
Veröffentlicht: (2022)
von: Fuhrer, Benjamin, et al.
Veröffentlicht: (2022)
Improving Token-Based World Models with Parallel Observation Prediction
von: Cohen, Lior, et al.
Veröffentlicht: (2024)
von: Cohen, Lior, et al.
Veröffentlicht: (2024)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
SQT -- std $Q$-target
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
von: Cohen, Lior, et al.
Veröffentlicht: (2025)
von: Cohen, Lior, et al.
Veröffentlicht: (2025)
Accelerating Vehicle Routing via AI-Initialized Genetic Algorithms
von: Greenberg, Ido, et al.
Veröffentlicht: (2025)
von: Greenberg, Ido, et al.
Veröffentlicht: (2025)
VLM-Guided Experience Replay
von: Sharony, Elad, et al.
Veröffentlicht: (2026)
von: Sharony, Elad, et al.
Veröffentlicht: (2026)
State Entropy Regularization for Robust Reinforcement Learning
von: Ashlag, Yonatan, et al.
Veröffentlicht: (2025)
von: Ashlag, Yonatan, et al.
Veröffentlicht: (2025)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
Adversarial Bandit over Bandits: Hierarchical Bandits for Online Configuration Management
von: Avin, Chen, et al.
Veröffentlicht: (2025)
von: Avin, Chen, et al.
Veröffentlicht: (2025)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
von: Koren, Uri, et al.
Veröffentlicht: (2025)
von: Koren, Uri, et al.
Veröffentlicht: (2025)
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
Information-Theoretic Generalization Bounds for Sequential Decision Making
von: Futami, Futoshi, et al.
Veröffentlicht: (2026)
von: Futami, Futoshi, et al.
Veröffentlicht: (2026)
Learning Multiple Initial Solutions to Optimization Problems
von: Sharony, Elad, et al.
Veröffentlicht: (2024)
von: Sharony, Elad, et al.
Veröffentlicht: (2024)
Online Sequential Decision-Making with Unknown Delays
von: Wu, Ping, et al.
Veröffentlicht: (2024)
von: Wu, Ping, et al.
Veröffentlicht: (2024)
Reward Design for Justifiable Sequential Decision-Making
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
PriorZero: Bridging Language Priors and World Models for Decision Making
von: Xiong, Junyu, et al.
Veröffentlicht: (2026)
von: Xiong, Junyu, et al.
Veröffentlicht: (2026)
Sequential Decision Making with Expert Demonstrations under Unobserved Heterogeneity
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2024)
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On Bits and Bandits: Quantifying the Regret-Information Trade-off
von: Shufaro, Itai, et al.
Veröffentlicht: (2024) -
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024) -
Tree Search-Based Policy Optimization under Stochastic Execution Delay
von: Valensi, David, et al.
Veröffentlicht: (2024) -
Policy Gradient with Tree Expansion
von: Dalal, Gal, et al.
Veröffentlicht: (2023) -
MinMaxMin $Q$-learning
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)