MinMaxMin $Q$-learning
Fuente:
arXiv
Saved in:
| Main Authors: | Soffair, Nitsan, Mannor, Shie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conservative DDPG -- Pessimistic RL without Ensemble
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
SQT -- std $Q$-target
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Solving Collaborative Dec-POMDPs with Deep Reinforcement Learning Heuristics
by: Soffair, Nitsan
Published: (2022)
by: Soffair, Nitsan
Published: (2022)
Markov flow policy -- deep MC
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Smooth Min-Max Monotonic Networks
by: Igel, Christian
Published: (2023)
by: Igel, Christian
Published: (2023)
Optimizing Agent Collaboration through Heuristic Multi-Agent Planning
by: Soffair, Nitsan
Published: (2023)
by: Soffair, Nitsan
Published: (2023)
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
by: Perets, Binyamin, et al.
Published: (2026)
by: Perets, Binyamin, et al.
Published: (2026)
Bayesian Neural Networks: A Min-Max Game Framework
by: Hong, Junping, et al.
Published: (2023)
by: Hong, Junping, et al.
Published: (2023)
MinMax Recurrent Neural Cascades
by: Ronca, Alessandro
Published: (2026)
by: Ronca, Alessandro
Published: (2026)
Representation-Driven Reinforcement Learning
by: Nabati, Ofir, et al.
Published: (2023)
by: Nabati, Ofir, et al.
Published: (2023)
Sobolev Space Regularised Pre Density Models
by: Kozdoba, Mark, et al.
Published: (2023)
by: Kozdoba, Mark, et al.
Published: (2023)
Bayesian Optimization for Function-Valued Responses under Min-Max Criteria
by: Ahadi, Pouya, et al.
Published: (2025)
by: Ahadi, Pouya, et al.
Published: (2025)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Multi-Objective Reinforcement Learning with Max-Min Criterion: A Game-Theoretic Approach
by: Byeon, Woohyeon, et al.
Published: (2025)
by: Byeon, Woohyeon, et al.
Published: (2025)
Improving Token-Based World Models with Parallel Observation Prediction
by: Cohen, Lior, et al.
Published: (2024)
by: Cohen, Lior, et al.
Published: (2024)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
by: Valensi, David, et al.
Published: (2024)
by: Valensi, David, et al.
Published: (2024)
iMTSP: Solving Min-Max Multiple Traveling Salesman Problem with Imperative Learning
by: Guo, Yifan, et al.
Published: (2024)
by: Guo, Yifan, et al.
Published: (2024)
The Max-Min Formulation of Multi-Objective Reinforcement Learning: From Theory to a Model-Free Algorithm
by: Park, Giseung, et al.
Published: (2024)
by: Park, Giseung, et al.
Published: (2024)
ADV-0: Closed-Loop Min-Max Adversarial Training for Long-Tail Robustness in Autonomous Driving
by: Nie, Tong, et al.
Published: (2026)
by: Nie, Tong, et al.
Published: (2026)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
by: Cohen, Lior, et al.
Published: (2025)
by: Cohen, Lior, et al.
Published: (2025)
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026)
by: Sharony, Elad, et al.
Published: (2026)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
by: Koren, Uri, et al.
Published: (2025)
by: Koren, Uri, et al.
Published: (2025)
Min-Max Optimisation for Nonconvex-Nonconcave Functions Using a Random Zeroth-Order Extragradient Algorithm
by: Farzin, Amir Ali, et al.
Published: (2025)
by: Farzin, Amir Ali, et al.
Published: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2025)
by: Kumar, Navdeep, et al.
Published: (2025)
DPN: Decoupling Partition and Navigation for Neural Solvers of Min-max Vehicle Routing Problems
by: Zheng, Zhi, et al.
Published: (2024)
by: Zheng, Zhi, et al.
Published: (2024)
Efficient Neural Combinatorial Optimization Solver for the Min-max Heterogeneous Capacitated Vehicle Routing Problem
by: Wu, Xuan, et al.
Published: (2025)
by: Wu, Xuan, et al.
Published: (2025)
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
by: Cheng, Jie, et al.
Published: (2025)
by: Cheng, Jie, et al.
Published: (2025)
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
by: Lab, Mind, et al.
Published: (2026)
by: Lab, Mind, et al.
Published: (2026)
Learning Multiple Initial Solutions to Optimization Problems
by: Sharony, Elad, et al.
Published: (2024)
by: Sharony, Elad, et al.
Published: (2024)
Bandit Max-Min Fair Allocation
by: Harada, Tsubasa, et al.
Published: (2025)
by: Harada, Tsubasa, et al.
Published: (2025)
Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)
by: Park, Ryan, et al.
Published: (2024)
Implementing Reinforcement Learning Datacenter Congestion Control in NVIDIA NICs
by: Fuhrer, Benjamin, et al.
Published: (2022)
by: Fuhrer, Benjamin, et al.
Published: (2022)
Min-$k$ Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics
by: Ding, Yuanhao, et al.
Published: (2026)
by: Ding, Yuanhao, et al.
Published: (2026)
Retain-Neutral Surrogates for Min-Max Unlearning
by: Cai, Junhao, et al.
Published: (2026)
by: Cai, Junhao, et al.
Published: (2026)
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Representative Action Selection for Large Action Space: From Bandits to MDPs
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
Diffusion Stochastic Optimization for Min-Max Problems
by: Cai, Haoyuan, et al.
Published: (2024)
by: Cai, Haoyuan, et al.
Published: (2024)
Similar Items
-
Conservative DDPG -- Pessimistic RL without Ensemble
by: Soffair, Nitsan, et al.
Published: (2024) -
SQT -- std $Q$-target
by: Soffair, Nitsan, et al.
Published: (2024) -
Solving Collaborative Dec-POMDPs with Deep Reinforcement Learning Heuristics
by: Soffair, Nitsan
Published: (2022) -
Markov flow policy -- deep MC
by: Soffair, Nitsan, et al.
Published: (2024) -
Smooth Min-Max Monotonic Networks
by: Igel, Christian
Published: (2023)