UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Lingling, Xue, Yongfu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization
by: Fu, Lingling, et al.
Published: (2026)
by: Fu, Lingling, et al.
Published: (2026)
Hierarchical LoRA MoE for Efficient CTR Model Scaling
by: Zeng, Zhichen, et al.
Published: (2025)
by: Zeng, Zhichen, et al.
Published: (2025)
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022)
by: Aouali, Imad, et al.
Published: (2022)
RewardRank: Optimizing True Learning-to-Rank Utility
by: Bhatt, Gaurav, et al.
Published: (2025)
by: Bhatt, Gaurav, et al.
Published: (2025)
RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems
by: Zeng, Wenwen, et al.
Published: (2026)
by: Zeng, Wenwen, et al.
Published: (2026)
EBaReT: Expert-guided Bag Reward Transformer for Auto Bidding
by: Li, Kaiyuan, et al.
Published: (2025)
by: Li, Kaiyuan, et al.
Published: (2025)
Orchestrating Heterogeneous Experts: A Scalable MoE Framework with Anisotropy-Preserving Fusion
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
Exploring Test-time Scaling via Prediction Merging on Large-Scale Recommendation
by: Lyu, Fuyuan, et al.
Published: (2025)
by: Lyu, Fuyuan, et al.
Published: (2025)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL
by: Wang, Youting, et al.
Published: (2026)
by: Wang, Youting, et al.
Published: (2026)
Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Training
by: Xia, Tianle, et al.
Published: (2026)
by: Xia, Tianle, et al.
Published: (2026)
Mitigating Pooling Bias in E-commerce Search via False Negative Estimation
by: Wang, Xiaochen, et al.
Published: (2023)
by: Wang, Xiaochen, et al.
Published: (2023)
BiCoRec: Bias-Mitigated Context-Aware Sequential Recommendation Model
by: Muthivhi, Mufhumudzi, et al.
Published: (2025)
by: Muthivhi, Mufhumudzi, et al.
Published: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Temporal Information Retrieval via Time-Specifier Model Merging
by: Han, SeungYoon, et al.
Published: (2025)
by: Han, SeungYoon, et al.
Published: (2025)
Analyzing and Mitigating Repetitions in Trip Recommendation
by: Shu, Wenzheng, et al.
Published: (2025)
by: Shu, Wenzheng, et al.
Published: (2025)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start Recommendation
by: Liu, Jialin, et al.
Published: (2026)
by: Liu, Jialin, et al.
Published: (2026)
On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme Classification
by: Prakash, Jatin, et al.
Published: (2024)
by: Prakash, Jatin, et al.
Published: (2024)
Optimizing Novelty of Top-k Recommendations using Large Language Models and Reinforcement Learning
by: Sharma, Amit, et al.
Published: (2024)
by: Sharma, Amit, et al.
Published: (2024)
Debiasing Message Passing to Mitigate Popularity Bias in GNN-based Collaborative Filtering
by: Islam, Md Aminul, et al.
Published: (2026)
by: Islam, Md Aminul, et al.
Published: (2026)
LIME: Link-based user-item Interaction Modeling with decoupled xor attention for Efficient test time scaling
by: Jiang, Yunjiang, et al.
Published: (2025)
by: Jiang, Yunjiang, et al.
Published: (2025)
Linear-PAL: A Lightweight Ranker for Mitigating Shortcut Learning in Personalized, High-Bias Tabular Ranking
by: Pawar, Vipul Dinesh
Published: (2025)
by: Pawar, Vipul Dinesh
Published: (2025)
CF-KAN: Kolmogorov-Arnold Network-based Collaborative Filtering to Mitigate Catastrophic Forgetting in Recommender Systems
by: Park, Jin-Duk, et al.
Published: (2024)
by: Park, Jin-Duk, et al.
Published: (2024)
Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT
by: Saad-Falcon, Jon, et al.
Published: (2024)
by: Saad-Falcon, Jon, et al.
Published: (2024)
Mitigating Exposure Bias in Online Learning to Rank Recommendation: A Novel Reward Model for Cascading Bandits
by: Mansoury, Masoud, et al.
Published: (2024)
by: Mansoury, Masoud, et al.
Published: (2024)
A Language-Driven Framework for Improving Personalized Recommendations: Merging LLMs with Traditional Algorithms
by: Goldstein, Aaron, et al.
Published: (2025)
by: Goldstein, Aaron, et al.
Published: (2025)
ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging
by: Verma, Neha, et al.
Published: (2026)
by: Verma, Neha, et al.
Published: (2026)
Cross-attention Secretly Performs Orthogonal Alignment in Recommendation Models
by: Lee, Hyunin, et al.
Published: (2025)
by: Lee, Hyunin, et al.
Published: (2025)
Constructing a Question-Answering Simulator through the Distillation of LLMs
by: Liu, Haipeng, et al.
Published: (2025)
by: Liu, Haipeng, et al.
Published: (2025)
Metric-agnostic Learning-to-Rank via Boosting and Rank Approximation
by: Gomez, Camilo, et al.
Published: (2026)
by: Gomez, Camilo, et al.
Published: (2026)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Predict Click-Through Rates with Deep Interest Network Model in E-commerce Advertising
by: Zhou, Chang, et al.
Published: (2024)
by: Zhou, Chang, et al.
Published: (2024)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
by: Wang, Mengru, et al.
Published: (2025)
by: Wang, Mengru, et al.
Published: (2025)
Knowledge Distillation for Enhancing Walmart E-commerce Search Relevance Using Large Language Models
by: Shang, Hongwei, et al.
Published: (2025)
by: Shang, Hongwei, et al.
Published: (2025)
Joint Model Parameter Scaling and Universal-Domain Data Integration for E-commerce Search Ranking
by: Yu, Liren, et al.
Published: (2026)
by: Yu, Liren, et al.
Published: (2026)
Fine-Tuning Diffusion-Based Recommender Systems via Reinforcement Learning with Reward Function Optimization
by: Hou, Yu, et al.
Published: (2025)
by: Hou, Yu, et al.
Published: (2025)
Do Not Wait: Learning Re-Ranking Model Without User Feedback At Serving Time in E-Commerce
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
Similar Items
-
TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization
by: Fu, Lingling, et al.
Published: (2026) -
Hierarchical LoRA MoE for Efficient CTR Model Scaling
by: Zeng, Zhichen, et al.
Published: (2025) -
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022) -
RewardRank: Optimizing True Learning-to-Rank Utility
by: Bhatt, Gaurav, et al.
Published: (2025) -
RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems
by: Zeng, Wenwen, et al.
Published: (2026)