Cost and Reward Infused Metric Elicitation
Fuente:
arXiv
Saved in:
| Main Authors: | Bhateja, Chethan, O'Brien, Joseph, Hashmi, Afnaan, Prakash, Eva |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
Action-Dependent Optimality-Preserving Reward Shaping
by: Forbes, Grant C., et al.
Published: (2025)
by: Forbes, Grant C., et al.
Published: (2025)
Potential-Based Reward Shaping For Intrinsic Motivation
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
by: Quadros, André, et al.
Published: (2025)
by: Quadros, André, et al.
Published: (2025)
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
by: Li, Bangzheng, et al.
Published: (2024)
by: Li, Bangzheng, et al.
Published: (2024)
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
Versatile Ordering Network: An Attention-based Neural Network for Ordering Across Scales and Quality Metrics
by: Yu, Zehua, et al.
Published: (2024)
by: Yu, Zehua, et al.
Published: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
In-situ Autoguidance: Eliciting Self-Correction in Diffusion Models
by: Gu, Enhao, et al.
Published: (2025)
by: Gu, Enhao, et al.
Published: (2025)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Difference Rewards Policy Gradients
by: Castellini, Jacopo, et al.
Published: (2020)
by: Castellini, Jacopo, et al.
Published: (2020)
Deep Variational Inference Symbolic Regression
by: Butterworth, James, et al.
Published: (2026)
by: Butterworth, James, et al.
Published: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
by: Hisaki, Yukinari, et al.
Published: (2024)
by: Hisaki, Yukinari, et al.
Published: (2024)
Improving ML Training Data with Gold-Standard Quality Metrics
by: Barrett, Leslie, et al.
Published: (2025)
by: Barrett, Leslie, et al.
Published: (2025)
Decoding Rewards in Competitive Games: Inverse Game Theory with Entropy Regularization
by: Liao, Junyi, et al.
Published: (2026)
by: Liao, Junyi, et al.
Published: (2026)
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
by: Son, Daniel, et al.
Published: (2025)
by: Son, Daniel, et al.
Published: (2025)
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
by: Mongaras, Gabriel, et al.
Published: (2026)
by: Mongaras, Gabriel, et al.
Published: (2026)
I-GLIDE: Input Groups for Latent Health Indicators in Degradation Estimation
by: Thil, Lucas, et al.
Published: (2025)
by: Thil, Lucas, et al.
Published: (2025)
FluidWorld: Reaction-Diffusion Dynamics as a Predictive Substrate for World Models
by: Polly, Fabien
Published: (2026)
by: Polly, Fabien
Published: (2026)
FSC-Net: Fast-Slow Consolidation Networks for Continual Learning
by: Gorrim, Mohamed El
Published: (2025)
by: Gorrim, Mohamed El
Published: (2025)
TFMAdapter: Lightweight Instance-Level Adaptation of Foundation Models for Forecasting with Covariates
by: Dange, Afrin, et al.
Published: (2025)
by: Dange, Afrin, et al.
Published: (2025)
Downsized and Compromised?: Assessing the Faithfulness of Model Compression
by: Kamal, Moumita, et al.
Published: (2025)
by: Kamal, Moumita, et al.
Published: (2025)
Learning Stochastic Nonlinear Dynamics with Embedded Latent Transfer Operators
by: Ke, Naichang, et al.
Published: (2025)
by: Ke, Naichang, et al.
Published: (2025)
Intervening to Learn and Compose Causally Disentangled Representations
by: Markham, Alex, et al.
Published: (2025)
by: Markham, Alex, et al.
Published: (2025)
Multiple Token Divergence: Measuring and Steering In-Context Computation Density
by: Herrmann, Vincent, et al.
Published: (2025)
by: Herrmann, Vincent, et al.
Published: (2025)
Strengthening the Internal Adversarial Robustness in Lifted Neural Networks
by: Zach, Christopher
Published: (2025)
by: Zach, Christopher
Published: (2025)
Pruning Spurious Subgraphs for Graph Out-of-Distribution Generalization
by: Yao, Tianjun, et al.
Published: (2025)
by: Yao, Tianjun, et al.
Published: (2025)
Exploring Neural Granger Causality with xLSTMs: Unveiling Temporal Dependencies in Complex Data
by: Poonia, Harsh, et al.
Published: (2025)
by: Poonia, Harsh, et al.
Published: (2025)
AdamNX: An Adam improvement algorithm based on a novel exponential decay mechanism for the second-order moment estimate
by: Zhu, Meng, et al.
Published: (2025)
by: Zhu, Meng, et al.
Published: (2025)
Symbol-Temporal Consistency Self-supervised Learning for Robust Time Series Classification
by: Garcia, Kevin, et al.
Published: (2025)
by: Garcia, Kevin, et al.
Published: (2025)
Sparse Concept Anchoring for Interpretable and Controllable Neural Representations
by: Fraser, Sandy, et al.
Published: (2025)
by: Fraser, Sandy, et al.
Published: (2025)
A Comparative Analysis of Reinforcement Learning and Conventional Deep Learning Approaches for Bearing Fault Diagnosis
by: Çakır, Efe, et al.
Published: (2025)
by: Çakır, Efe, et al.
Published: (2025)
HGCN(O): A Self-Tuning GCN HyperModel Toolkit for Outcome Prediction in Event-Sequence Data
by: Wang, Fang, et al.
Published: (2025)
by: Wang, Fang, et al.
Published: (2025)
Local-Order Auxiliary Losses Can Improve Autoencoder Reconstruction
by: Dam, Harvey, et al.
Published: (2025)
by: Dam, Harvey, et al.
Published: (2025)
Hierarchical Reinforcement Learning with Targeted Causal Interventions
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
BOND: License to Train with Black-Box Functions
by: Clark, Andrew, et al.
Published: (2025)
by: Clark, Andrew, et al.
Published: (2025)
Similar Items
-
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025) -
Action-Dependent Optimality-Preserving Reward Shaping
by: Forbes, Grant C., et al.
Published: (2025) -
Potential-Based Reward Shaping For Intrinsic Motivation
by: Forbes, Grant C., et al.
Published: (2024) -
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
by: Quadros, André, et al.
Published: (2025) -
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
by: Li, Bangzheng, et al.
Published: (2024)