Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Lingkai, Satish, Anagha, Jiang, Hezi, Kangaslahti, Akseli, Ma, Andrew, Chen, Wenbo, Song, Mingxiao, Xu, Lily, Tambe, Milind |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Policy-Embedded Graph Expansion: Networked HIV Testing with Diffusion-Driven Network Samples
by: Kangaslahti, Akseli, et al.
Published: (2026)
by: Kangaslahti, Akseli, et al.
Published: (2026)
Generative AI Against Poaching: Latent Composite Flow Matching for Wildlife Conservation
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Network-Based Interventions for HIV Prevention via Cascade-Aware Suppression of Transmission
by: Kangaslahti, Akseli, et al.
Published: (2026)
by: Kangaslahti, Akseli, et al.
Published: (2026)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
by: Wang, Haichuan, et al.
Published: (2026)
by: Wang, Haichuan, et al.
Published: (2026)
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
by: Verma, Shresth, et al.
Published: (2024)
by: Verma, Shresth, et al.
Published: (2024)
Generative AI for Social Impact
by: Kong, Lingkai, et al.
Published: (2026)
by: Kong, Lingkai, et al.
Published: (2026)
Robust Optimization with Diffusion Models for Green Security
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement Learning
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
Reinforcement learning with combinatorial actions for coupled restless bandits
by: Xu, Lily, et al.
Published: (2025)
by: Xu, Lily, et al.
Published: (2025)
LLM-based Agent Simulation for Maternal Health Interventions: Uncertainty Estimation and Decision-focused Evaluation
by: Martinson, Sarah, et al.
Published: (2025)
by: Martinson, Sarah, et al.
Published: (2025)
Dynamic Targeting of Satellite Observations Using Supplemental Geostationary Satellite Data and Hierarchical Planning
by: Kangaslahti, Akseli, et al.
Published: (2026)
by: Kangaslahti, Akseli, et al.
Published: (2026)
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
Combining Diverse Information for Coordinated Action: Stochastic Bandit Algorithms for Heterogeneous Agents
by: Gordon, Lucia, et al.
Published: (2024)
by: Gordon, Lucia, et al.
Published: (2024)
Dual-Mandate Patrols: Multi-Armed Bandits for Green Security
by: Xu, Lily, et al.
Published: (2020)
by: Xu, Lily, et al.
Published: (2020)
Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits
by: Liang, Biyonka, et al.
Published: (2024)
by: Liang, Biyonka, et al.
Published: (2024)
VORTEX: Aligning Task Utility and Human Preferences through LLM-Guided Reward Shaping
by: Xiong, Guojun, et al.
Published: (2025)
by: Xiong, Guojun, et al.
Published: (2025)
Learning-Based Planning for Improving Science Return of Earth Observation Satellites
by: Breitfeld, Abigail, et al.
Published: (2025)
by: Breitfeld, Abigail, et al.
Published: (2025)
Artificial Replay: A Meta-Algorithm for Harnessing Historical Data in Bandits
by: Banerjee, Siddhartha, et al.
Published: (2022)
by: Banerjee, Siddhartha, et al.
Published: (2022)
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025)
by: Tec, Mauricio, et al.
Published: (2025)
Flow-Based Policy for Online Reinforcement Learning
by: Lv, Lei, et al.
Published: (2025)
by: Lv, Lei, et al.
Published: (2025)
On Diffusion Models for Multi-Agent Partial Observability: Shared Attractors, Error Bounds, and Composite Flow
by: Wang, Tonghan, et al.
Published: (2024)
by: Wang, Tonghan, et al.
Published: (2024)
Leaving the Nest: Going Beyond Local Loss Functions for Predict-Then-Optimize
by: Shah, Sanket, et al.
Published: (2023)
by: Shah, Sanket, et al.
Published: (2023)
Contrasting local and global modeling with machine learning and satellite data: A case study estimating tree canopy height in African savannas
by: Rolf, Esther, et al.
Published: (2024)
by: Rolf, Esther, et al.
Published: (2024)
Learning to Persuade a Biased Receiver
by: Pan, Yuqi, et al.
Published: (2026)
by: Pan, Yuqi, et al.
Published: (2026)
Many Preferences, Few Policies: Towards Scalable Language Model Personalization
by: Kim, Cheol Woo, et al.
Published: (2026)
by: Kim, Cheol Woo, et al.
Published: (2026)
Analyzing Cost-Sensitive Surrogate Losses via $\mathcal{H}$-calibration
by: Shah, Sanket, et al.
Published: (2025)
by: Shah, Sanket, et al.
Published: (2025)
Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement Learning
by: Yu, Dongjie, et al.
Published: (2026)
by: Yu, Dongjie, et al.
Published: (2026)
Escape Sensing Games: Detection-vs-Evasion in Security Applications
by: Boehmer, Niclas, et al.
Published: (2024)
by: Boehmer, Niclas, et al.
Published: (2024)
Efficient Public Health Intervention Planning Using Decomposition-Based Decision-Focused Learning
by: Shah, Sanket, et al.
Published: (2024)
by: Shah, Sanket, et al.
Published: (2024)
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
by: Yang, Xinyi, et al.
Published: (2025)
by: Yang, Xinyi, et al.
Published: (2025)
Embeddings for Preferences, Not Semantics
by: Blair, Carter, et al.
Published: (2026)
by: Blair, Carter, et al.
Published: (2026)
Constrained Latent Action Policies for Model-Based Offline Reinforcement Learning
by: Alles, Marvin, et al.
Published: (2024)
by: Alles, Marvin, et al.
Published: (2024)
Quantitative Convergences of Lie Group Momentum Optimizers
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
Convergence of Kinetic Langevin Monte Carlo on Lie groups
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
The Bandit Whisperer: Communication Learning for Restless Bandits
by: Zhao, Yunfan, et al.
Published: (2024)
by: Zhao, Yunfan, et al.
Published: (2024)
Fairness for Workers Who Pull the Arms: An Index Based Policy for Allocation of Restless Bandit Tasks
by: Biswas, Arpita, et al.
Published: (2023)
by: Biswas, Arpita, et al.
Published: (2023)
Towards Foundation-model-based Multiagent System to Accelerate AI for Social Impact
by: Zhao, Yunfan, et al.
Published: (2024)
by: Zhao, Yunfan, et al.
Published: (2024)
Where to Intervene: Action Selection in Deep Reinforcement Learning
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation
by: Songwei, Wu, et al.
Published: (2026)
by: Songwei, Wu, et al.
Published: (2026)
Similar Items
-
Policy-Embedded Graph Expansion: Networked HIV Testing with Diffusion-Driven Network Samples
by: Kangaslahti, Akseli, et al.
Published: (2026) -
Generative AI Against Poaching: Latent Composite Flow Matching for Wildlife Conservation
by: Kong, Lingkai, et al.
Published: (2025) -
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
by: Kong, Lingkai, et al.
Published: (2025) -
Network-Based Interventions for HIV Prevention via Cascade-Aware Suppression of Transmission
by: Kangaslahti, Akseli, et al.
Published: (2026) -
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
by: Wang, Haichuan, et al.
Published: (2026)