K-Score: Kalman Filter as a Principled Alternative to Reward Normalization in Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Xia, Zixuan, Li, Quanxi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning
por: Li, Yuxuan, et al.
Publicado: (2025)
por: Li, Yuxuan, et al.
Publicado: (2025)
Rewarding the Journey, Not Just the Destination: A Composite Path and Answer Self-Scoring Reward Mechanism for Test-Time Reinforcement Learning
por: Xing, Jingyu, et al.
Publicado: (2025)
por: Xing, Jingyu, et al.
Publicado: (2025)
Kalman Filter for Online Classification of Non-Stationary Data
por: Titsias, Michalis K., et al.
Publicado: (2023)
por: Titsias, Michalis K., et al.
Publicado: (2023)
Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise
por: Long, Bo, et al.
Publicado: (2026)
por: Long, Bo, et al.
Publicado: (2026)
Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics
por: Huml, JR, et al.
Publicado: (2026)
por: Huml, JR, et al.
Publicado: (2026)
Reinforcement Learning with Symbolic Reward Machines
por: Krug, Thomas, et al.
Publicado: (2026)
por: Krug, Thomas, et al.
Publicado: (2026)
Reinforcement Learning with Exogenous States and Rewards
por: Trimponias, George, et al.
Publicado: (2023)
por: Trimponias, George, et al.
Publicado: (2023)
Reinforcement Learning with Stochastic Reward Machines
por: Corazza, Jan, et al.
Publicado: (2025)
por: Corazza, Jan, et al.
Publicado: (2025)
Offline Reinforcement Learning with Imputed Rewards
por: Romeo, Carlo, et al.
Publicado: (2024)
por: Romeo, Carlo, et al.
Publicado: (2024)
2048: Reinforcement Learning in a Delayed Reward Environment
por: Saligram, Prady, et al.
Publicado: (2025)
por: Saligram, Prady, et al.
Publicado: (2025)
Is Optimal Transport Necessary for Inverse Reinforcement Learning?
por: Dong, Zixuan, et al.
Publicado: (2025)
por: Dong, Zixuan, et al.
Publicado: (2025)
Beyond Rewards in Reinforcement Learning for Cyber Defence
por: Bates, Elizabeth, et al.
Publicado: (2026)
por: Bates, Elizabeth, et al.
Publicado: (2026)
RLSR: Reinforcement Learning from Self Reward
por: Simonds, Toby, et al.
Publicado: (2025)
por: Simonds, Toby, et al.
Publicado: (2025)
Efficient Reinforcement Learning in Probabilistic Reward Machines
por: Lin, Xiaofeng, et al.
Publicado: (2024)
por: Lin, Xiaofeng, et al.
Publicado: (2024)
Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
por: Chaudhari, Shreyas, et al.
Publicado: (2025)
por: Chaudhari, Shreyas, et al.
Publicado: (2025)
Learning a Diffusion Model Policy from Rewards via Q-Score Matching
por: Psenka, Michael, et al.
Publicado: (2023)
por: Psenka, Michael, et al.
Publicado: (2023)
iPINNER: An Iterative Physics-Informed Neural Network with Ensemble Kalman Filter
por: Lu, Binghang, et al.
Publicado: (2025)
por: Lu, Binghang, et al.
Publicado: (2025)
Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
por: Abdi, Hossein, et al.
Publicado: (2025)
por: Abdi, Hossein, et al.
Publicado: (2025)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
por: Ishihara, Yu, et al.
Publicado: (2025)
por: Ishihara, Yu, et al.
Publicado: (2025)
Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
por: Xu, Peiran, et al.
Publicado: (2025)
por: Xu, Peiran, et al.
Publicado: (2025)
Shaping Sparse Rewards in Reinforcement Learning: A Semi-supervised Approach
por: Li, Wenyun, et al.
Publicado: (2025)
por: Li, Wenyun, et al.
Publicado: (2025)
Adaptive Correlation-Weighted Intrinsic Rewards for Reinforcement Learning
por: Nguyen, Viet Bac, et al.
Publicado: (2026)
por: Nguyen, Viet Bac, et al.
Publicado: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
por: Lu, Xiaodong, et al.
Publicado: (2026)
por: Lu, Xiaodong, et al.
Publicado: (2026)
Reward Models in Deep Reinforcement Learning: A Survey
por: Yu, Rui, et al.
Publicado: (2025)
por: Yu, Rui, et al.
Publicado: (2025)
Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?
por: Omori, Yumi, et al.
Publicado: (2025)
por: Omori, Yumi, et al.
Publicado: (2025)
FM-EAC: Feature Model-based Enhanced Actor-Critic for Multi-Task Control in Dynamic Environments
por: Zhou, Quanxi, et al.
Publicado: (2025)
por: Zhou, Quanxi, et al.
Publicado: (2025)
Normality-Guided Distributional Reinforcement Learning for Continuous Control
por: Byun, Ju-Seung, et al.
Publicado: (2022)
por: Byun, Ju-Seung, et al.
Publicado: (2022)
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
por: Zhang, Feng, et al.
Publicado: (2026)
por: Zhang, Feng, et al.
Publicado: (2026)
Reinfier and Reintrainer: Verification and Interpretation-Driven Safe Deep Reinforcement Learning Frameworks
por: Yang, Zixuan, et al.
Publicado: (2024)
por: Yang, Zixuan, et al.
Publicado: (2024)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
por: Pukdee, Rattana, et al.
Publicado: (2026)
por: Pukdee, Rattana, et al.
Publicado: (2026)
CHARM: Calibrating Reward Models With Chatbot Arena Scores
por: Zhu, Xiao, et al.
Publicado: (2025)
por: Zhu, Xiao, et al.
Publicado: (2025)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
por: Zhang, Zijing, et al.
Publicado: (2025)
por: Zhang, Zijing, et al.
Publicado: (2025)
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
por: Zamboni, Riccardo, et al.
Publicado: (2025)
por: Zamboni, Riccardo, et al.
Publicado: (2025)
RewardAnything: Generalizable Principle-Following Reward Models
por: Yu, Zhuohao, et al.
Publicado: (2025)
por: Yu, Zhuohao, et al.
Publicado: (2025)
Contrastive Representation for Data Filtering in Cross-Domain Offline Reinforcement Learning
por: Wen, Xiaoyu, et al.
Publicado: (2024)
por: Wen, Xiaoyu, et al.
Publicado: (2024)
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
por: Li, Mengdi, et al.
Publicado: (2025)
por: Li, Mengdi, et al.
Publicado: (2025)
Reward-Punishment Reinforcement Learning with Maximum Entropy
por: Wang, Jiexin, et al.
Publicado: (2024)
por: Wang, Jiexin, et al.
Publicado: (2024)
Reinforcement Learning with Reward Machines for Sleep Control in Mobile Networks
por: Levina, Kristina, et al.
Publicado: (2026)
por: Levina, Kristina, et al.
Publicado: (2026)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
por: Kar, Avik, et al.
Publicado: (2024)
por: Kar, Avik, et al.
Publicado: (2024)
Decision-Focused Model-based Reinforcement Learning for Reward Transfer
por: Sharma, Abhishek, et al.
Publicado: (2023)
por: Sharma, Abhishek, et al.
Publicado: (2023)
Ejemplares similares
-
TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning
por: Li, Yuxuan, et al.
Publicado: (2025) -
Rewarding the Journey, Not Just the Destination: A Composite Path and Answer Self-Scoring Reward Mechanism for Test-Time Reinforcement Learning
por: Xing, Jingyu, et al.
Publicado: (2025) -
Kalman Filter for Online Classification of Non-Stationary Data
por: Titsias, Michalis K., et al.
Publicado: (2023) -
Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise
por: Long, Bo, et al.
Publicado: (2026) -
Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics
por: Huml, JR, et al.
Publicado: (2026)