Optimal Transport for LLM Reward Modeling from Noisy Preference
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Licheng, Yang, Haochen, Li, Haoxuan, Lu, Yunsheng, Tong, Yongqi, Wang, Yinuo, Wang, Shijian, Chu, Zhixuan, Shen, Lei, Lu, Yuan, Wang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Reward Modeling for Large Language Models via Causal Decomposition
by: Lu, Yunsheng, et al.
Published: (2026)
by: Lu, Yunsheng, et al.
Published: (2026)
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary
by: Pan, Licheng, et al.
Published: (2025)
by: Pan, Licheng, et al.
Published: (2025)
Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast Models
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Deep Time-series Forecasting Needs Kernelized Moment Balancing
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
by: Han, Haochen, et al.
Published: (2024)
by: Han, Haochen, et al.
Published: (2024)
Deep Autocorrelation Modeling for Time-Series Forecasting: Progress and Prospects
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Lowest Span Confidence: A Zero-Shot Metric for Efficient and Black-Box Hallucination Detection in LLMs
by: Qiao, Yitong, et al.
Published: (2026)
by: Qiao, Yitong, et al.
Published: (2026)
Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
by: Wang, Shuqiang, et al.
Published: (2026)
by: Wang, Shuqiang, et al.
Published: (2026)
Toward Optimal Statistical Inference in Noisy Linear Quadratic Reinforcement Learning over a Finite Horizon
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
by: Li, Zhuo, et al.
Published: (2025)
by: Li, Zhuo, et al.
Published: (2025)
Debiased Recommendation with Noisy Feedback
by: Li, Haoxuan, et al.
Published: (2024)
by: Li, Haoxuan, et al.
Published: (2024)
Time-o1: Time-Series Forecasting Needs Transformed Label Alignment
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Learning Ordinal Probabilistic Reward from Preferences
by: Chen, Longze, et al.
Published: (2026)
by: Chen, Longze, et al.
Published: (2026)
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
by: Peng, Hao, et al.
Published: (2025)
by: Peng, Hao, et al.
Published: (2025)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
IterSIMP-σ: Evaluating LLM-Assisted Spatial Interventions in Stress-Aware Topology Optimization
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
Analyzing and Improving Diffusion Models for Time-Series Data Imputation: A Proximal Recursion Perspective
by: Chen, Zhichao, et al.
Published: (2026)
by: Chen, Zhichao, et al.
Published: (2026)
Optimal Transport for Treatment Effect Estimation
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
AutoSiMP: Autonomous Topology Optimization from Natural Language via LLM-Driven Problem Configuration and Adaptive Solver Control
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
Understanding Zero-shot Rare Word Recognition Improvements Through LLM Integration
by: Wang, Haoxuan
Published: (2025)
by: Wang, Haoxuan
Published: (2025)
Reward Training Wheels: Adaptive Auxiliary Rewards for Robotics Reinforcement Learning
by: Wang, Linji, et al.
Published: (2025)
by: Wang, Linji, et al.
Published: (2025)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
by: Pan, Zhixuan, et al.
Published: (2025)
by: Pan, Zhixuan, et al.
Published: (2025)
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
by: Cai, Hongru, et al.
Published: (2026)
by: Cai, Hongru, et al.
Published: (2026)
BaseReward: A Strong Baseline for Multimodal Reward Model
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Self-Ensemble Post Learning for Noisy Domain Generalization
by: Lu, Wang, et al.
Published: (2025)
by: Lu, Wang, et al.
Published: (2025)
PrefMoE: Robust Preference Modeling with Mixture-of-Experts Reward Learning
by: Yuan, Ziqin, et al.
Published: (2026)
by: Yuan, Ziqin, et al.
Published: (2026)
Integrating Optimal Transport and Structural Inference Models for GRN Inference from Single-cell Data
by: Tong, Tsz Pan, et al.
Published: (2024)
by: Tong, Tsz Pan, et al.
Published: (2024)
RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
by: Song, Yushuai, et al.
Published: (2026)
by: Song, Yushuai, et al.
Published: (2026)
Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
by: Yang, Wenhao, et al.
Published: (2024)
by: Yang, Wenhao, et al.
Published: (2024)
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
Learning Pareto-Optimal Rewards from Noisy Preferences: A Framework for Multi-Objective Inverse Reinforcement Learning
by: Cherukuri, Kalyan, et al.
Published: (2025)
by: Cherukuri, Kalyan, et al.
Published: (2025)
PLOT: Enhancing Preference Learning via Optimal Transport
by: Zhu, Liang, et al.
Published: (2026)
by: Zhu, Liang, et al.
Published: (2026)
Large Language Models as Optimization Controllers: Adaptive Continuation for SIMP Topology Optimization
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
OmniGAIA: Towards Native Omni-Modal AI Agents
by: Li, Xiaoxi, et al.
Published: (2026)
by: Li, Xiaoxi, et al.
Published: (2026)
Similar Items
-
Robust Reward Modeling for Large Language Models via Causal Decomposition
by: Lu, Yunsheng, et al.
Published: (2026) -
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
by: Wang, Hao, et al.
Published: (2026) -
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026) -
Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary
by: Pan, Licheng, et al.
Published: (2025) -
Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast Models
by: Wang, Hao, et al.
Published: (2025)