Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Shakerinava, Mehran, Ravanbakhsh, Siamak, Oberman, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Role of Symmetry in Optimizing Overparameterized Networks
by: Sareen, Kusha, et al.
Published: (2026)
by: Sareen, Kusha, et al.
Published: (2026)
The Expressive Limits of Diagonal SSMs for State-Tracking
by: Shakerinava, Mehran, et al.
Published: (2026)
by: Shakerinava, Mehran, et al.
Published: (2026)
Weight-Sharing Regularization
by: Shakerinava, Mehran, et al.
Published: (2023)
by: Shakerinava, Mehran, et al.
Published: (2023)
Learning to Reach Goals via Diffusion
by: Jain, Vineet, et al.
Published: (2023)
by: Jain, Vineet, et al.
Published: (2023)
Scaling Laws and Symmetry, Evidence from Neural Force Fields
by: Ngo, Khang, et al.
Published: (2025)
by: Ngo, Khang, et al.
Published: (2025)
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
by: Jain, Vineet, et al.
Published: (2025)
by: Jain, Vineet, et al.
Published: (2025)
On Diffusion Modeling for Anomaly Detection
by: Livernoche, Victor, et al.
Published: (2023)
by: Livernoche, Victor, et al.
Published: (2023)
Inverting Data Transformations via Diffusion Sampling
by: Kim, Jinwoo, et al.
Published: (2026)
by: Kim, Jinwoo, et al.
Published: (2026)
Energy Loss Functions for Physical Systems
by: Kaba, Sékou-Oumar, et al.
Published: (2025)
by: Kaba, Sékou-Oumar, et al.
Published: (2025)
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
by: Kim, Kihyun, et al.
Published: (2025)
by: Kim, Kihyun, et al.
Published: (2025)
Efficient Dynamics Modeling in Interactive Environments with Koopman Theory
by: Mondal, Arnab Kumar, et al.
Published: (2023)
by: Mondal, Arnab Kumar, et al.
Published: (2023)
Parity Requires Unified Input Dependence and Negative Eigenvalues in SSMs
by: Khavari, Behnoush, et al.
Published: (2025)
by: Khavari, Behnoush, et al.
Published: (2025)
Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits
by: Xue, Bo, et al.
Published: (2025)
by: Xue, Bo, et al.
Published: (2025)
Local Inconsistency Resolution: The Interplay between Attention and Control in Probabilistic Models
by: Richardson, Oliver E., et al.
Published: (2026)
by: Richardson, Oliver E., et al.
Published: (2026)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
Beyond MSE: Improving Precipitation Nowcasting with Multi-Quantile Regression
by: van Nieuwkoop, Gijs, et al.
Published: (2026)
by: van Nieuwkoop, Gijs, et al.
Published: (2026)
Thresholded Lexicographic Ordered Multiobjective Reinforcement Learning
by: Tercan, Alperen, et al.
Published: (2024)
by: Tercan, Alperen, et al.
Published: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026)
by: Muni, Aneri, et al.
Published: (2026)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
by: Ye, Ziyi, et al.
Published: (2024)
by: Ye, Ziyi, et al.
Published: (2024)
Beyond Scalar Rewards: Distributional Reinforcement Learning with Preordered Objectives for Safe and Reliable Autonomous Driving
by: Abouelazm, Ahmed, et al.
Published: (2026)
by: Abouelazm, Ahmed, et al.
Published: (2026)
LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
by: Qiu, Ruiyu, et al.
Published: (2025)
by: Qiu, Ruiyu, et al.
Published: (2025)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
by: Afsharrad, Amirhossein, et al.
Published: (2026)
by: Afsharrad, Amirhossein, et al.
Published: (2026)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Axiomatization of Gradient Smoothing in Neural Networks
by: Zhou, Linjiang, et al.
Published: (2024)
by: Zhou, Linjiang, et al.
Published: (2024)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Beyond Rewards in Reinforcement Learning for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2026)
by: Bates, Elizabeth, et al.
Published: (2026)
Beyond Distribution Sharpening: The Importance of Task Rewards
by: Mittal, Sarthak, et al.
Published: (2026)
by: Mittal, Sarthak, et al.
Published: (2026)
A Diffusion-Contrastive Graph Neural Network with Virtual Nodes for Wind Nowcasting in Unobserved Regions
by: Shi, Jie, et al.
Published: (2026)
by: Shi, Jie, et al.
Published: (2026)
EEG-MFTNet: An Enhanced EEGNet Architecture with Multi-Scale Temporal Convolutions and Transformer Fusion for Cross-Session Motor Imagery Decoding
by: Andrikopoulos, Panagiotis, et al.
Published: (2026)
by: Andrikopoulos, Panagiotis, et al.
Published: (2026)
SmaAT-QMix-UNet: A Parameter-Efficient Vector-Quantized UNet for Precipitation Nowcasting
by: Stavrou, Nikolas, et al.
Published: (2026)
by: Stavrou, Nikolas, et al.
Published: (2026)
Progressive Inference-Time Annealing of Diffusion Models for Sampling from Boltzmann Densities
by: Akhound-Sadegh, Tara, et al.
Published: (2025)
by: Akhound-Sadegh, Tara, et al.
Published: (2025)
STARC: A General Framework For Quantifying Differences Between Reward Functions
by: Skalse, Joar, et al.
Published: (2023)
by: Skalse, Joar, et al.
Published: (2023)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Multi-Armed Sampling Problem and the End of Exploration
by: Pedramfar, Mohammad, et al.
Published: (2025)
by: Pedramfar, Mohammad, et al.
Published: (2025)
Thinking Beyond Visibility: A Near-Optimal Policy Framework for Locally Interdependent Multi-Agent MDPs
by: DeWeese, Alex, et al.
Published: (2025)
by: DeWeese, Alex, et al.
Published: (2025)
Lexicographic optimization-based approaches to learning a representative model for multi-criteria sorting with non-monotonic criteria
by: Zhang, Zhen, et al.
Published: (2024)
by: Zhang, Zhen, et al.
Published: (2024)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
Similar Items
-
The Role of Symmetry in Optimizing Overparameterized Networks
by: Sareen, Kusha, et al.
Published: (2026) -
The Expressive Limits of Diagonal SSMs for State-Tracking
by: Shakerinava, Mehran, et al.
Published: (2026) -
Weight-Sharing Regularization
by: Shakerinava, Mehran, et al.
Published: (2023) -
Learning to Reach Goals via Diffusion
by: Jain, Vineet, et al.
Published: (2023) -
Scaling Laws and Symmetry, Evidence from Neural Force Fields
by: Ngo, Khang, et al.
Published: (2025)