Bi-Level Policy Optimization with Nyström Hypergradients
Fuente:
arXiv
Saved in:
| Main Authors: | Prakash, Arjun, He, Naicheng, Goktas, Denizalp, Makar-Limanov, Jacob, Greenwald, Amy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convex-Concave Zero-sum Markov Stackelberg Games
by: Goktas, Denizalp, et al.
Published: (2024)
by: Goktas, Denizalp, et al.
Published: (2024)
Efficient Inverse Multiagent Learning
by: Goktas, Denizalp, et al.
Published: (2025)
by: Goktas, Denizalp, et al.
Published: (2025)
Tractable General Equilibrium
by: Goktas, Denizalp, et al.
Published: (2025)
by: Goktas, Denizalp, et al.
Published: (2025)
Banzhaf Power in Hierarchical Voting Games
by: Randolph, John, et al.
Published: (2025)
by: Randolph, John, et al.
Published: (2025)
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026)
by: Kudo, Mikoto, et al.
Published: (2026)
Tâtonnement in Homothetic Fisher Markets
by: Goktas, Denizalp, et al.
Published: (2023)
by: Goktas, Denizalp, et al.
Published: (2023)
Infinite Horizon Markov Economies
by: Goktas, Denizalp, et al.
Published: (2025)
by: Goktas, Denizalp, et al.
Published: (2025)
NePPO: Near-Potential Policy Optimization for General-Sum Multi-Agent Reinforcement Learning
by: Kalanther, Addison, et al.
Published: (2026)
by: Kalanther, Addison, et al.
Published: (2026)
Policy Aggregation
by: Alamdari, Parand A., et al.
Published: (2024)
by: Alamdari, Parand A., et al.
Published: (2024)
A Resilience Framework for Bi-Criteria Combinatorial Optimization with Bandit Feedback
by: Aggarwal, Vaneet, et al.
Published: (2025)
by: Aggarwal, Vaneet, et al.
Published: (2025)
Proximal Policy Optimization with Evolutionary Mutations
by: Czworkowski, Casimir, et al.
Published: (2026)
by: Czworkowski, Casimir, et al.
Published: (2026)
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
by: Zhang, Yuheng, et al.
Published: (2024)
by: Zhang, Yuheng, et al.
Published: (2024)
Learning in Markov Games with Adaptive Adversaries: Policy Regret, Fundamental Barriers, and Efficient Algorithms
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
by: Liu, Mingyang, et al.
Published: (2024)
by: Liu, Mingyang, et al.
Published: (2024)
Human Choice Prediction in Language-based Persuasion Games: Simulation-based Off-Policy Evaluation
by: Shapira, Eilam, et al.
Published: (2023)
by: Shapira, Eilam, et al.
Published: (2023)
Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models
by: Hennes, Daniel, et al.
Published: (2026)
by: Hennes, Daniel, et al.
Published: (2026)
Policy Optimization finds Nash Equilibrium in Regularized General-Sum LQ Games
by: Zaman, Muhammad Aneeq uz, et al.
Published: (2024)
by: Zaman, Muhammad Aneeq uz, et al.
Published: (2024)
Optimizing Hard-to-Place Kidney Allocation: A Machine Learning Approach to Center Ranking
by: Berry, Sean, et al.
Published: (2024)
by: Berry, Sean, et al.
Published: (2024)
Model-Based RL for Mean-Field Games is not Statistically Harder than Single-Agent RL
by: Huang, Jiawei, et al.
Published: (2024)
by: Huang, Jiawei, et al.
Published: (2024)
Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles
by: Lian, Jiesong, et al.
Published: (2024)
by: Lian, Jiesong, et al.
Published: (2024)
Reach Measurement, Optimization and Frequency Capping In Targeted Online Advertising Under k-Anonymity
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
ELA: Exploited Level Augmentation for Offline Learning in Zero-Sum Games
by: Lei, Shiqi, et al.
Published: (2024)
by: Lei, Shiqi, et al.
Published: (2024)
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
by: Wolf, Will
Published: (2025)
by: Wolf, Will
Published: (2025)
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
by: Hornig, Benedikt, et al.
Published: (2026)
by: Hornig, Benedikt, et al.
Published: (2026)
Efficient Ensemble Selection from Binary and Pairwise Feedback
by: Neoh, Tzeh Yuan, et al.
Published: (2026)
by: Neoh, Tzeh Yuan, et al.
Published: (2026)
Learning and Collusion in Multi-unit Auctions
by: Brânzei, Simina, et al.
Published: (2023)
by: Brânzei, Simina, et al.
Published: (2023)
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
by: Liu, Mingyang, et al.
Published: (2024)
by: Liu, Mingyang, et al.
Published: (2024)
Self-optimization in distributed manufacturing systems using Modular State-based Stackelberg Games
by: Yuwono, Steve, et al.
Published: (2024)
by: Yuwono, Steve, et al.
Published: (2024)
Generative Social Choice: The Next Generation
by: Boehmer, Niclas, et al.
Published: (2025)
by: Boehmer, Niclas, et al.
Published: (2025)
LLM-Auction: Generative Auction towards LLM-Native Advertising
by: Zhao, Chujie, et al.
Published: (2025)
by: Zhao, Chujie, et al.
Published: (2025)
Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
by: Ai, Rui, et al.
Published: (2025)
by: Ai, Rui, et al.
Published: (2025)
From Leiden to Pleasure Island: The Constant Potts Model for Community Detection as a Hedonic Game
by: Felipe, Lucas Lopes, et al.
Published: (2025)
by: Felipe, Lucas Lopes, et al.
Published: (2025)
GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series
by: Kim, Jinwoong, et al.
Published: (2026)
by: Kim, Jinwoong, et al.
Published: (2026)
Robust Deep Monte Carlo Counterfactual Regret Minimization: Addressing Theoretical Risks in Neural Fictitious Self-Play
by: Jaafari, Zakaria El
Published: (2025)
by: Jaafari, Zakaria El
Published: (2025)
Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning
by: Waugh, Justin
Published: (2026)
by: Waugh, Justin
Published: (2026)
Machine Learning-Powered Course Allocation
by: Soumalias, Ermis, et al.
Published: (2022)
by: Soumalias, Ermis, et al.
Published: (2022)
PerfectDou: Dominating DouDizhu with Perfect Information Distillation
by: Yang, Guan, et al.
Published: (2022)
by: Yang, Guan, et al.
Published: (2022)
Proportional Fairness in Non-Centroid Clustering
by: Caragiannis, Ioannis, et al.
Published: (2024)
by: Caragiannis, Ioannis, et al.
Published: (2024)
Similar Items
-
Convex-Concave Zero-sum Markov Stackelberg Games
by: Goktas, Denizalp, et al.
Published: (2024) -
Efficient Inverse Multiagent Learning
by: Goktas, Denizalp, et al.
Published: (2025) -
Tractable General Equilibrium
by: Goktas, Denizalp, et al.
Published: (2025) -
Banzhaf Power in Hierarchical Voting Games
by: Randolph, John, et al.
Published: (2025) -
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026)