StepScorer: Accelerating Reinforcement Learning with Step-wise Scoring and Psychological Regret Modeling
Fuente:
arXiv
Guardado en:
| Autor principal: | Xu, Zhe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
por: Yousaf, Iqra
Publicado: (2024)
por: Yousaf, Iqra
Publicado: (2024)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
por: Tiwari, Dhruv
Publicado: (2025)
por: Tiwari, Dhruv
Publicado: (2025)
Fractional Policy Gradients: Reinforcement Learning with Long-Term Memory
por: Pawar, Urvi, et al.
Publicado: (2025)
por: Pawar, Urvi, et al.
Publicado: (2025)
From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning
por: Tomashevskiy, Timofey
Publicado: (2026)
por: Tomashevskiy, Timofey
Publicado: (2026)
Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts
por: Chapman, James, et al.
Publicado: (2025)
por: Chapman, James, et al.
Publicado: (2025)
FlowRL: Flow-Augmented Few-Shot Reinforcement Learning for Semi-Structured Sensor Data
por: Pivezhandi, Mohammad, et al.
Publicado: (2024)
por: Pivezhandi, Mohammad, et al.
Publicado: (2024)
A Multidisciplinary Approach to Telegram Data Analysis
por: Varbanov, Velizar, et al.
Publicado: (2024)
por: Varbanov, Velizar, et al.
Publicado: (2024)
Regret-Aware Policy Optimization: Environment-Level Memory for Replay Suppression under Delayed Harm
por: Hiremath, Prakul Sunil
Publicado: (2026)
por: Hiremath, Prakul Sunil
Publicado: (2026)
Predicting and improving test-time scaling laws via reward tail-guided search
por: Li, Muheng, et al.
Publicado: (2026)
por: Li, Muheng, et al.
Publicado: (2026)
On the Limits of Learned Importance Scoring for KV Cache Compression
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
An Aircraft Upset Recovery System with Reinforcement Learning
por: Demir, Mahir, et al.
Publicado: (2026)
por: Demir, Mahir, et al.
Publicado: (2026)
Safe Reinforcement Learning with Preference-based Constraint Inference
por: Li, Chenglin, et al.
Publicado: (2026)
por: Li, Chenglin, et al.
Publicado: (2026)
CORE: Towards Scalable and Efficient Causal Discovery with Reinforcement Learning
por: Sauter, Andreas W. M., et al.
Publicado: (2024)
por: Sauter, Andreas W. M., et al.
Publicado: (2024)
GIRL: Generative Imagination Reinforcement Learning via Information-Theoretic Hallucination Control
por: Hiremath, Prakul Sunil
Publicado: (2026)
por: Hiremath, Prakul Sunil
Publicado: (2026)
Deep Reinforcement Learning for Day-to-day Dynamic Tolling in Tradable Credit Schemes
por: Wu, Xiaoyi, et al.
Publicado: (2025)
por: Wu, Xiaoyi, et al.
Publicado: (2025)
Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems
por: Nuzhin, Egor E., et al.
Publicado: (2024)
por: Nuzhin, Egor E., et al.
Publicado: (2024)
Score-informed Neural Operator for Enhancing Ordering-based Causal Discovery
por: Kang, Jiyeon, et al.
Publicado: (2025)
por: Kang, Jiyeon, et al.
Publicado: (2025)
Joint Combinatorial Node Selection and Resource Allocations in the Lightning Network using Attention-based Reinforcement Learning
por: Salahshour, Mahdi, et al.
Publicado: (2024)
por: Salahshour, Mahdi, et al.
Publicado: (2024)
Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance
por: Silue, Bram, et al.
Publicado: (2025)
por: Silue, Bram, et al.
Publicado: (2025)
Diffusion-MPC in Discrete Domains: Feasibility Constraints, Horizon Effects, and Critic Alignment: Case study with Tetris
por: Wang, Haochuan Kevin
Publicado: (2026)
por: Wang, Haochuan Kevin
Publicado: (2026)
The Final-Stage Bottleneck: A Systematic Dissection of the R-Learner for Network Causal Inference
por: Sairam, S, et al.
Publicado: (2025)
por: Sairam, S, et al.
Publicado: (2025)
What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
por: Zhang, Xinyu
Publicado: (2026)
por: Zhang, Xinyu
Publicado: (2026)
Factored Diffusion Policies:Compositionally Generalized Robot Control with a Single Score Network
por: Mitra, Sayan, et al.
Publicado: (2026)
por: Mitra, Sayan, et al.
Publicado: (2026)
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning
por: Hu, Yuelin, et al.
Publicado: (2026)
por: Hu, Yuelin, et al.
Publicado: (2026)
Are We Winning the Wrong Game? Revisiting Evaluation Practices for Long-Term Time Series Forecasting
por: Phungtua-eng, Thanapol, et al.
Publicado: (2026)
por: Phungtua-eng, Thanapol, et al.
Publicado: (2026)
Machine Learning Based Path Planning for Improved Rover Navigation (Pre-Print Version)
por: Abcouwer, Neil, et al.
Publicado: (2020)
por: Abcouwer, Neil, et al.
Publicado: (2020)
AGWM: Affordance-Grounded World Models for Environments with Compositional Prerequisites
por: Zhang, Qinshi, et al.
Publicado: (2026)
por: Zhang, Qinshi, et al.
Publicado: (2026)
A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
por: Kujur, Arahan
Publicado: (2026)
por: Kujur, Arahan
Publicado: (2026)
Working Paper: Active Causal Structure Learning with Latent Variables: Towards Learning to Detour in Autonomous Robots
por: Riscos, Pablo de los, et al.
Publicado: (2024)
por: Riscos, Pablo de los, et al.
Publicado: (2024)
An Automatic Ground Collision Avoidance System with Reinforcement Learning
por: Sevgili, Seyyid Osman, et al.
Publicado: (2026)
por: Sevgili, Seyyid Osman, et al.
Publicado: (2026)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
por: Vincze, Mátyás, et al.
Publicado: (2024)
por: Vincze, Mátyás, et al.
Publicado: (2024)
Adaptable Hindsight Experience Replay for Search-Based Learning
por: Vazaios, Alexandros, et al.
Publicado: (2025)
por: Vazaios, Alexandros, et al.
Publicado: (2025)
Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
por: Wang, Zizhao, et al.
Publicado: (2024)
por: Wang, Zizhao, et al.
Publicado: (2024)
Graph Neural Networks are Heuristics
por: Min, Yimeng, et al.
Publicado: (2026)
por: Min, Yimeng, et al.
Publicado: (2026)
Differentiable Symbolic Planning: A Neural Architecture for Constraint Reasoning with Learned Feasibility
por: Oruganti, Venkatakrishna Reddy
Publicado: (2026)
por: Oruganti, Venkatakrishna Reddy
Publicado: (2026)
Selective Progress-Aware Querying for Human-in-the-Loop Reinforcement Learning
por: Muraleedharan, Anujith, et al.
Publicado: (2025)
por: Muraleedharan, Anujith, et al.
Publicado: (2025)
Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
por: Xu, Jiexi
Publicado: (2025)
por: Xu, Jiexi
Publicado: (2025)
Social Interpretable Reinforcement Learning
por: Custode, Leonardo Lucio, et al.
Publicado: (2024)
por: Custode, Leonardo Lucio, et al.
Publicado: (2024)
Reinforcement Learning for Dynamic Workflow Optimization in CI/CD Pipelines
por: Soni, Aniket Abhishek, et al.
Publicado: (2026)
por: Soni, Aniket Abhishek, et al.
Publicado: (2026)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
por: Fang, Qitong, et al.
Publicado: (2026)
por: Fang, Qitong, et al.
Publicado: (2026)
Ejemplares similares
-
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
por: Yousaf, Iqra
Publicado: (2024) -
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
por: Tiwari, Dhruv
Publicado: (2025) -
Fractional Policy Gradients: Reinforcement Learning with Long-Term Memory
por: Pawar, Urvi, et al.
Publicado: (2025) -
From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning
por: Tomashevskiy, Timofey
Publicado: (2026) -
Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts
por: Chapman, James, et al.
Publicado: (2025)