Saved in:
| Main Authors: | Xiong, Guojun, Tambe, Milind |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.16399 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Health Facility Location in Ethiopia: Leveraging LLMs to Integrate Expert Knowledge into Algorithmic Planning
by: Trabelsi, Yohai, et al.
Published: (2026)
by: Trabelsi, Yohai, et al.
Published: (2026)
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025)
by: Tec, Mauricio, et al.
Published: (2025)
Embeddings for Preferences, Not Semantics
by: Blair, Carter, et al.
Published: (2026)
by: Blair, Carter, et al.
Published: (2026)
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
by: Verma, Shresth, et al.
Published: (2024)
by: Verma, Shresth, et al.
Published: (2024)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
by: Wang, Haichuan, et al.
Published: (2026)
by: Wang, Haichuan, et al.
Published: (2026)
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
Optimizing Vital Sign Monitoring in Resource-Constrained Maternal Care: An RL-Based Restless Bandit Approach
by: Boehmer, Niclas, et al.
Published: (2024)
by: Boehmer, Niclas, et al.
Published: (2024)
Combining Diverse Information for Coordinated Action: Stochastic Bandit Algorithms for Heterogeneous Agents
by: Gordon, Lucia, et al.
Published: (2024)
by: Gordon, Lucia, et al.
Published: (2024)
LLM-based Agent Simulation for Maternal Health Interventions: Uncertainty Estimation and Decision-focused Evaluation
by: Martinson, Sarah, et al.
Published: (2025)
by: Martinson, Sarah, et al.
Published: (2025)
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
Towards Foundation-model-based Multiagent System to Accelerate AI for Social Impact
by: Zhao, Yunfan, et al.
Published: (2024)
by: Zhao, Yunfan, et al.
Published: (2024)
LLM Active Alignment: A Nash Equilibrium Perspective
by: Wang, Tonghan, et al.
Published: (2026)
by: Wang, Tonghan, et al.
Published: (2026)
Many Preferences, Few Policies: Towards Scalable Language Model Personalization
by: Kim, Cheol Woo, et al.
Published: (2026)
by: Kim, Cheol Woo, et al.
Published: (2026)
Leaving the Nest: Going Beyond Local Loss Functions for Predict-Then-Optimize
by: Shah, Sanket, et al.
Published: (2023)
by: Shah, Sanket, et al.
Published: (2023)
Efficient Public Health Intervention Planning Using Decomposition-Based Decision-Focused Learning
by: Shah, Sanket, et al.
Published: (2024)
by: Shah, Sanket, et al.
Published: (2024)
Fairness for Workers Who Pull the Arms: An Index Based Policy for Allocation of Restless Bandit Tasks
by: Biswas, Arpita, et al.
Published: (2023)
by: Biswas, Arpita, et al.
Published: (2023)
A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health
by: Behari, Nikhil, et al.
Published: (2024)
by: Behari, Nikhil, et al.
Published: (2024)
Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
by: Ai, Rui, et al.
Published: (2025)
by: Ai, Rui, et al.
Published: (2025)
On Sequential Fault-Intolerant Process Planning
by: Kaczmarczyk, Andrzej, et al.
Published: (2025)
by: Kaczmarczyk, Andrzej, et al.
Published: (2025)
Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective
by: Kim, Cheol Woo, et al.
Published: (2026)
by: Kim, Cheol Woo, et al.
Published: (2026)
Reinforcement learning with combinatorial actions for coupled restless bandits
by: Xu, Lily, et al.
Published: (2025)
by: Xu, Lily, et al.
Published: (2025)
Online Allocation with Unknown Shared Supply
by: Neoh, Tzeh Yuan, et al.
Published: (2026)
by: Neoh, Tzeh Yuan, et al.
Published: (2026)
Contrasting local and global modeling with machine learning and satellite data: A case study estimating tree canopy height in African savannas
by: Rolf, Esther, et al.
Published: (2024)
by: Rolf, Esther, et al.
Published: (2024)
Aligning Crowd Feedback via Distributional Preference Reward Modeling
by: Li, Dexun, et al.
Published: (2024)
by: Li, Dexun, et al.
Published: (2024)
Decisions and Deployment: The Five-Year SAHELI Project (2020-2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health
by: Verma, Shresth, et al.
Published: (2026)
by: Verma, Shresth, et al.
Published: (2026)
OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings
by: Zaregarizi, Shadmehr, et al.
Published: (2026)
by: Zaregarizi, Shadmehr, et al.
Published: (2026)
From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge
by: Wu, Xiefeng
Published: (2024)
by: Wu, Xiefeng
Published: (2024)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
by: Lin, Wenze, et al.
Published: (2026)
by: Lin, Wenze, et al.
Published: (2026)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
by: Bhambri, Siddhant, et al.
Published: (2023)
by: Bhambri, Siddhant, et al.
Published: (2023)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
by: Wu, Boyuan
Published: (2025)
by: Wu, Boyuan
Published: (2025)
Beyond Listenership: AI-Predicted Interventions Drive Improvements in Maternal Health Behaviours
by: Dasgupta, Arpan, et al.
Published: (2025)
by: Dasgupta, Arpan, et al.
Published: (2025)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025)
by: Wu, Keming, et al.
Published: (2025)
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
by: Shankar, Shreya, et al.
Published: (2024)
by: Shankar, Shreya, et al.
Published: (2024)
Learning to Align Human Code Preferences
by: Yin, Xin, et al.
Published: (2025)
by: Yin, Xin, et al.
Published: (2025)
LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping Robots
by: Han, Dongge, et al.
Published: (2024)
by: Han, Dongge, et al.
Published: (2024)
MTRec: Learning to Align with User Preferences via Mental Reward Models
by: Zhao, Mengchen, et al.
Published: (2025)
by: Zhao, Mengchen, et al.
Published: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
by: Jain, Gauri, et al.
Published: (2024)
by: Jain, Gauri, et al.
Published: (2024)
Capturing Individual Human Preferences with Reward Features
by: Barreto, André, et al.
Published: (2025)
by: Barreto, André, et al.
Published: (2025)
Similar Items
-
Health Facility Location in Ethiopia: Leveraging LLMs to Integrate Expert Knowledge into Algorithmic Planning
by: Trabelsi, Yohai, et al.
Published: (2026) -
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
by: Kong, Lingkai, et al.
Published: (2025) -
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025) -
Embeddings for Preferences, Not Semantics
by: Blair, Carter, et al.
Published: (2026) -
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
by: Verma, Shresth, et al.
Published: (2024)