Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Haichuan, Lin, Tao, Kong, Lingkai, Li, Ce, Jiang, Hezi, Tambe, Milind |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
by: Verma, Shresth, et al.
Published: (2024)
by: Verma, Shresth, et al.
Published: (2024)
Generative AI Against Poaching: Latent Composite Flow Matching for Wildlife Conservation
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
by: Rodriguez-Diaz, Paula, et al.
Published: (2024)
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025)
by: Tec, Mauricio, et al.
Published: (2025)
Harnesses for Inference-Time Alignment over Execution Trajectories
by: Wang, Boyuan, et al.
Published: (2026)
by: Wang, Boyuan, et al.
Published: (2026)
VORTEX: Aligning Task Utility and Human Preferences through LLM-Guided Reward Shaping
by: Xiong, Guojun, et al.
Published: (2025)
by: Xiong, Guojun, et al.
Published: (2025)
Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective
by: Kim, Cheol Woo, et al.
Published: (2026)
by: Kim, Cheol Woo, et al.
Published: (2026)
Combining Diverse Information for Coordinated Action: Stochastic Bandit Algorithms for Heterogeneous Agents
by: Gordon, Lucia, et al.
Published: (2024)
by: Gordon, Lucia, et al.
Published: (2024)
Leaving the Nest: Going Beyond Local Loss Functions for Predict-Then-Optimize
by: Shah, Sanket, et al.
Published: (2023)
by: Shah, Sanket, et al.
Published: (2023)
Efficient Public Health Intervention Planning Using Decomposition-Based Decision-Focused Learning
by: Shah, Sanket, et al.
Published: (2024)
by: Shah, Sanket, et al.
Published: (2024)
Reinforcement learning with combinatorial actions for coupled restless bandits
by: Xu, Lily, et al.
Published: (2025)
by: Xu, Lily, et al.
Published: (2025)
Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions
by: Kong, Lingkai, et al.
Published: (2026)
by: Kong, Lingkai, et al.
Published: (2026)
Decisions and Deployment: The Five-Year SAHELI Project (2020-2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health
by: Verma, Shresth, et al.
Published: (2026)
by: Verma, Shresth, et al.
Published: (2026)
Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review
by: Uehara, Masatoshi, et al.
Published: (2025)
by: Uehara, Masatoshi, et al.
Published: (2025)
Bandits with Preference Feedback: A Stackelberg Game Perspective
by: Pásztor, Barna, et al.
Published: (2024)
by: Pásztor, Barna, et al.
Published: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Learning Graph Structures and Uncertainty for Accurate and Calibrated Time-series Forecasting
by: Kamarthi, Harshavardhan, et al.
Published: (2024)
by: Kamarthi, Harshavardhan, et al.
Published: (2024)
LLM-based Agent Simulation for Maternal Health Interventions: Uncertainty Estimation and Decision-focused Evaluation
by: Martinson, Sarah, et al.
Published: (2025)
by: Martinson, Sarah, et al.
Published: (2025)
PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
by: Lin, Baijiong, et al.
Published: (2025)
by: Lin, Baijiong, et al.
Published: (2025)
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
Trivialized Momentum Facilitates Diffusion Generative Modeling on Lie Groups
by: Zhu, Yuchen, et al.
Published: (2024)
by: Zhu, Yuchen, et al.
Published: (2024)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
by: Jain, Gauri, et al.
Published: (2024)
by: Jain, Gauri, et al.
Published: (2024)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
by: Lin, Wenze, et al.
Published: (2026)
by: Lin, Wenze, et al.
Published: (2026)
Stackelberg Coupling of Online Representation Learning and Reinforcement Learning
by: Martinez, Fernando, et al.
Published: (2025)
by: Martinez, Fernando, et al.
Published: (2025)
Online Prompt Pricing based on Combinatorial Multi-Armed Bandit and Hierarchical Stackelberg Game
by: Li, Meiling, et al.
Published: (2024)
by: Li, Meiling, et al.
Published: (2024)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Contrasting local and global modeling with machine learning and satellite data: A case study estimating tree canopy height in African savannas
by: Rolf, Esther, et al.
Published: (2024)
by: Rolf, Esther, et al.
Published: (2024)
Learnable Chernoff Baselines for Inference-Time Alignment
by: Madhow, Sunil, et al.
Published: (2026)
by: Madhow, Sunil, et al.
Published: (2026)
Inference-Time Scaling for Generalist Reward Modeling
by: Liu, Zijun, et al.
Published: (2025)
by: Liu, Zijun, et al.
Published: (2025)
Evaluating the Effectiveness of Index-Based Treatment Allocation
by: Boehmer, Niclas, et al.
Published: (2024)
by: Boehmer, Niclas, et al.
Published: (2024)
Robust Optimization with Diffusion Models for Green Security
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Reward-free Alignment for Conflicting Objectives
by: Chen, Peter, et al.
Published: (2026)
by: Chen, Peter, et al.
Published: (2026)
Dynamic Search for Inference-Time Alignment in Diffusion Models
by: Li, Xiner, et al.
Published: (2025)
by: Li, Xiner, et al.
Published: (2025)
LLM Active Alignment: A Nash Equilibrium Perspective
by: Wang, Tonghan, et al.
Published: (2026)
by: Wang, Tonghan, et al.
Published: (2026)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Two Birds with One Stone: Enhancing Uncertainty Quantification and Interpretability with Graph Functional Neural Process
by: Kong, Lingkai, et al.
Published: (2025)
by: Kong, Lingkai, et al.
Published: (2025)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
by: Holmes, Ian, et al.
Published: (2025)
by: Holmes, Ian, et al.
Published: (2025)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
by: Huang, Audrey, et al.
Published: (2025)
by: Huang, Audrey, et al.
Published: (2025)
Similar Items
-
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
by: Kong, Lingkai, et al.
Published: (2025) -
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
by: Verma, Shresth, et al.
Published: (2024) -
Generative AI Against Poaching: Latent Composite Flow Matching for Wildlife Conservation
by: Kong, Lingkai, et al.
Published: (2025) -
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
by: Rodriguez-Diaz, Paula, et al.
Published: (2024) -
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025)