Debiasing Online Preference Learning via Preference Feature Preservation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Dongyoung, Yoon, Jinsung, Shin, Jinwoo, Kim, Jaehyung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Training-free LLM Verification via Recycling Few-shot Examples
by: Lee, Dongseok, et al.
Published: (2025)
by: Lee, Dongseok, et al.
Published: (2025)
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
Verifier-free Test-Time Sampling for Vision Language Action Models
by: Jang, Suhyeok, et al.
Published: (2025)
by: Jang, Suhyeok, et al.
Published: (2025)
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
by: Kim, Dongyoung, et al.
Published: (2025)
by: Kim, Dongyoung, et al.
Published: (2025)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Enhancing Instruction Following of LLMs via Activation Steering with Dynamic Rejection
by: Kang, Minjae, et al.
Published: (2026)
by: Kang, Minjae, et al.
Published: (2026)
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
by: Kim, Gihoon, et al.
Published: (2026)
by: Kim, Gihoon, et al.
Published: (2026)
PassREfinder-FL: Privacy-Preserving Credential Stuffing Risk Prediction via Graph-Based Federated Learning for Representing Password Reuse between Websites
by: Kim, Jaehan, et al.
Published: (2025)
by: Kim, Jaehan, et al.
Published: (2025)
Tabular Transfer Learning via Prompting LLMs
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
Efficient LLM Collaboration via Planning
by: Lee, Byeongchan, et al.
Published: (2025)
by: Lee, Byeongchan, et al.
Published: (2025)
InterPol: De-anonymizing LM Arena via Interpolated Preference Learning
by: Cho, Minsung, et al.
Published: (2026)
by: Cho, Minsung, et al.
Published: (2026)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
by: Song, Woomin, et al.
Published: (2024)
by: Song, Woomin, et al.
Published: (2024)
Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
by: Nam, Yunhun, et al.
Published: (2025)
by: Nam, Yunhun, et al.
Published: (2025)
From Efficiency to Equity: Measuring Fairness in Preference Learning
by: Gowaikar, Shreeyash, et al.
Published: (2024)
by: Gowaikar, Shreeyash, et al.
Published: (2024)
Evaluating Feature Dependent Noise in Preference-based Reinforcement Learning
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
Hindsight Preference Learning for Offline Preference-based Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2024)
by: Gao, Chen-Xiao, et al.
Published: (2024)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
by: Yoon, Hee Suk, et al.
Published: (2025)
by: Yoon, Hee Suk, et al.
Published: (2025)
Forget and Explain: Transparent Verification of GNN Unlearning
by: Ahsan, Imran, et al.
Published: (2025)
by: Ahsan, Imran, et al.
Published: (2025)
Shortcut Features as Top Eigenfunctions of NTK: A Linear Neural Network Case and More
by: Lim, Jinwoo, et al.
Published: (2026)
by: Lim, Jinwoo, et al.
Published: (2026)
Physics-Guided Geometric Diffusion for Macro Placement Generation
by: Yoon, Jongho, et al.
Published: (2026)
by: Yoon, Jongho, et al.
Published: (2026)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
by: Kim, Dongyoung, et al.
Published: (2026)
by: Kim, Dongyoung, et al.
Published: (2026)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
by: Kwak, Minseo, et al.
Published: (2026)
by: Kwak, Minseo, et al.
Published: (2026)
Few-shot Personalization of LLMs with Mis-aligned Responses
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Capturing Individual Human Preferences with Reward Features
by: Barreto, André, et al.
Published: (2025)
by: Barreto, André, et al.
Published: (2025)
EdgeGFL: Rethinking Edge Information in Graph Feature Preference Learning
by: Zhuo, Shengda, et al.
Published: (2025)
by: Zhuo, Shengda, et al.
Published: (2025)
Structural Reasoning Improves Molecular Understanding of LLM
by: Jang, Yunhui, et al.
Published: (2024)
by: Jang, Yunhui, et al.
Published: (2024)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
by: Driss, Brahim, et al.
Published: (2025)
by: Driss, Brahim, et al.
Published: (2025)
T-POP: Test-Time Personalization with Online Preference Feedback
by: Qu, Zikun, et al.
Published: (2025)
by: Qu, Zikun, et al.
Published: (2025)
Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
by: Kim, Jung-hun, et al.
Published: (2017)
by: Kim, Jung-hun, et al.
Published: (2017)
Trust Region Q Adjoint Matching
by: Dong, Yonghoon, et al.
Published: (2026)
by: Dong, Yonghoon, et al.
Published: (2026)
Querying Easily Flip-flopped Samples for Deep Active Learning
by: Cho, Seong Jin, et al.
Published: (2024)
by: Cho, Seong Jin, et al.
Published: (2024)
Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional Anisotropy
by: Jeong, Wooseong, et al.
Published: (2026)
by: Jeong, Wooseong, et al.
Published: (2026)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning
by: Diwan, Nirav, et al.
Published: (2025)
by: Diwan, Nirav, et al.
Published: (2025)
Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Similar Items
-
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024) -
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024) -
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023) -
Training-free LLM Verification via Recycling Few-shot Examples
by: Lee, Dongseok, et al.
Published: (2025) -
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
by: Nam, Jaehyun, et al.
Published: (2024)