PRISM: Parallel Reward Integration with Symmetry for MORL
Fuente:
arXiv
Saved in:
| Main Authors: | van der Knaap, Finn, Qian, Kejiang, Xu, Zheng, He, Fengxiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
by: Tang, Kenton, et al.
Published: (2026)
by: Tang, Kenton, et al.
Published: (2026)
Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
by: Zuo, Qian, et al.
Published: (2025)
by: Zuo, Qian, et al.
Published: (2025)
Learning Pareto-Optimal Pandemic Intervention Policies with MORL
by: Chen, Marian, et al.
Published: (2025)
by: Chen, Marian, et al.
Published: (2025)
PA2D-MORL: Pareto Ascent Directional Decomposition based Multi-Objective Reinforcement Learning
by: Hu, Tianmeng, et al.
Published: (2026)
by: Hu, Tianmeng, et al.
Published: (2026)
Rationality Measurement and Theory for Reinforcement Learning Agents
by: Qian, Kejiang, et al.
Published: (2026)
by: Qian, Kejiang, et al.
Published: (2026)
Integrating LTL Constraints into PPO for Safe Reinforcement Learning
by: Zhang, Maifang, et al.
Published: (2026)
by: Zhang, Maifang, et al.
Published: (2026)
PRISM: Mitigating EHR Data Sparsity via Learning from Missing Feature Calibrated Prototype Patient Representations
by: Zhu, Yinghao, et al.
Published: (2023)
by: Zhu, Yinghao, et al.
Published: (2023)
PRISM: Purified Representation and Integrated Semantic Modeling for Generative Sequential Recommendation
by: Fang, Dengzhao, et al.
Published: (2026)
by: Fang, Dengzhao, et al.
Published: (2026)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models
by: Wang, Xuliang, et al.
Published: (2026)
by: Wang, Xuliang, et al.
Published: (2026)
When a Reinforcement Learning Agent Encounters Unknown Unknowns
by: Zhu, Juntian, et al.
Published: (2025)
by: Zhu, Juntian, et al.
Published: (2025)
PRISM: Structured Optimization via Anisotropic Spectral Shaping
by: Yang, Yujie
Published: (2026)
by: Yang, Yujie
Published: (2026)
In-Context Reward Adaptation for Robust Preference Modeling
by: Sun, Zhenyu, et al.
Published: (2026)
by: Sun, Zhenyu, et al.
Published: (2026)
DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks
by: Shu, Aijie, et al.
Published: (2026)
by: Shu, Aijie, et al.
Published: (2026)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
by: Sun, Hao, et al.
Published: (2023)
by: Sun, Hao, et al.
Published: (2023)
XAI for In-hospital Mortality Prediction via Multimodal ICU Data
by: Li, Xingqiao, et al.
Published: (2023)
by: Li, Xingqiao, et al.
Published: (2023)
PRISM-CTG: A Foundation Model for Cardiotocography Analysis with Multi-View SSL
by: Wong, Sheng, et al.
Published: (2026)
by: Wong, Sheng, et al.
Published: (2026)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
by: Zhao, Yang, et al.
Published: (2026)
by: Zhao, Yang, et al.
Published: (2026)
"I May Not Have Articulated Myself Clearly": Diagnosing Dynamic Instability in LLM Reasoning at Inference Time
by: Chen, Jinkun, et al.
Published: (2026)
by: Chen, Jinkun, et al.
Published: (2026)
PRISM: Exploring Heterogeneous Pretrained EEG Foundation Model Transfer to Clinical Differential Diagnosis
by: Lahiri, Jeet Bandhu, et al.
Published: (2026)
by: Lahiri, Jeet Bandhu, et al.
Published: (2026)
PRISM: Lightweight Multivariate Time-Series Classification through Symmetric Multi-Resolution Convolutional Layers
by: Zucchi, Federico, et al.
Published: (2025)
by: Zucchi, Federico, et al.
Published: (2025)
Causal Deep Learning
by: Berrevoets, Jeroen, et al.
Published: (2023)
by: Berrevoets, Jeroen, et al.
Published: (2023)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Rectifying Shortcut Behaviors in Preference-based Reward Learning
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
by: Rafailov, Rafael, et al.
Published: (2023)
by: Rafailov, Rafael, et al.
Published: (2023)
NUTS, NARS, and Speech
by: van der Sluis, D.
Published: (2024)
by: van der Sluis, D.
Published: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search
by: Mou, Zhiyu, et al.
Published: (2025)
by: Mou, Zhiyu, et al.
Published: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Exploiting Symmetry in Dynamics for Model-Based Reinforcement Learning with Asymmetric Rewards
by: Sonmez, Yasin, et al.
Published: (2024)
by: Sonmez, Yasin, et al.
Published: (2024)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
by: Kopf, Laura, et al.
Published: (2025)
by: Kopf, Laura, et al.
Published: (2025)
PRISM: Privacy-preserving Inference System with Homomorphic Encryption and Modular Activation
by: Elkhatib, Zeinab, et al.
Published: (2025)
by: Elkhatib, Zeinab, et al.
Published: (2025)
Generalized Parallel Scaling with Interdependent Generations
by: Dong, Harry, et al.
Published: (2025)
by: Dong, Harry, et al.
Published: (2025)
Value Flows
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
by: Cho, Minjae, et al.
Published: (2026)
by: Cho, Minjae, et al.
Published: (2026)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
by: Holmes, Ian, et al.
Published: (2025)
by: Holmes, Ian, et al.
Published: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Similar Items
-
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
by: Tang, Kenton, et al.
Published: (2026) -
Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
by: Zuo, Qian, et al.
Published: (2025) -
Learning Pareto-Optimal Pandemic Intervention Policies with MORL
by: Chen, Marian, et al.
Published: (2025) -
PA2D-MORL: Pareto Ascent Directional Decomposition based Multi-Objective Reinforcement Learning
by: Hu, Tianmeng, et al.
Published: (2026) -
Rationality Measurement and Theory for Reinforcement Learning Agents
by: Qian, Kejiang, et al.
Published: (2026)