Preference Optimization by Estimating the Ratio of the Data Distribution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Yeongmin, Bae, Heesun, Na, Byeonghu, Moon, Il-Chul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reward-based Input Construction for Cross-document Relation Extraction
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
AMiD: Knowledge Distillation for LLMs with $α$-mixture Assistant Distribution
von: Shin, Donghyeok, et al.
Veröffentlicht: (2025)
von: Shin, Donghyeok, et al.
Veröffentlicht: (2025)
Distillation of Large Language Models via Concrete Score Matching
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
von: Na, Byeonghu, et al.
Veröffentlicht: (2026)
von: Na, Byeonghu, et al.
Veröffentlicht: (2026)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
von: Kim, Yeongmin, et al.
Veröffentlicht: (2026)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2026)
Prompt-Based Safety Guidance Is Ineffective for Unlearned Text-to-Image Diffusion Models
von: Shin, Jiwoo, et al.
Veröffentlicht: (2025)
von: Shin, Jiwoo, et al.
Veröffentlicht: (2025)
Diffusion Bridge AutoEncoders for Unsupervised Representation Learning
von: Kim, Yeongmin, et al.
Veröffentlicht: (2024)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2024)
Diffusion Rejection Sampling
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
Dirichlet-based Per-Sample Weighting by Transition Matrix for Noisy Label Learning
von: Bae, HeeSun, et al.
Veröffentlicht: (2024)
von: Bae, HeeSun, et al.
Veröffentlicht: (2024)
Training Unbiased Diffusion Models From Biased Dataset
von: Kim, Yeongmin, et al.
Veröffentlicht: (2024)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2024)
Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
von: Na, Byeonghu, et al.
Veröffentlicht: (2025)
von: Na, Byeonghu, et al.
Veröffentlicht: (2025)
Selective Preference Optimization via Token-Level Reward Function Estimation
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
On the Role of Preference Variance in Preference Optimization
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Label-Noise Robust Diffusion Models
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
von: Na, Byeonghu, et al.
Veröffentlicht: (2025)
von: Na, Byeonghu, et al.
Veröffentlicht: (2025)
Geometric-Averaged Preference Optimization for Soft Preference Labels
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Adaptive Guidance for Retrieval-Augmented Masked Diffusion Models
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
Filtered Direct Preference Optimization
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2024)
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2024)
Direct Preference Optimization with an Offset
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
von: Jian, Chengtao, et al.
Veröffentlicht: (2025)
von: Jian, Chengtao, et al.
Veröffentlicht: (2025)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
Towards Improved Preference Optimization Pipeline: from Data Generation to Budget-Controlled Regularization
von: Chen, Zhuotong, et al.
Veröffentlicht: (2024)
von: Chen, Zhuotong, et al.
Veröffentlicht: (2024)
Entropy Controllable Direct Preference Optimization
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
Orthogonal Finetuning for Direct Preference Optimization
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
VPO: Leveraging the Number of Votes in Preference Optimization
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
TSO: Self-Training with Scaled Preference Optimization
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Efficient and Scalable Estimation of Tool Representations in Vector Space
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
von: Yang, Junming, et al.
Veröffentlicht: (2025)
von: Yang, Junming, et al.
Veröffentlicht: (2025)
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2026)
von: Takahashi, Hiroshi, et al.
Veröffentlicht: (2026)
ROPO: Robust Preference Optimization for Large Language Models
von: Liang, Xize, et al.
Veröffentlicht: (2024)
von: Liang, Xize, et al.
Veröffentlicht: (2024)
Accelerated Preference Optimization for Large Language Model Alignment
von: He, Jiafan, et al.
Veröffentlicht: (2024)
von: He, Jiafan, et al.
Veröffentlicht: (2024)
Self-Play Preference Optimization for Language Model Alignment
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reward-based Input Construction for Cross-document Relation Extraction
von: Na, Byeonghu, et al.
Veröffentlicht: (2024) -
AMiD: Knowledge Distillation for LLMs with $α$-mixture Assistant Distribution
von: Shin, Donghyeok, et al.
Veröffentlicht: (2025) -
Distillation of Large Language Models via Concrete Score Matching
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025) -
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
von: Na, Byeonghu, et al.
Veröffentlicht: (2026) -
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
von: Kim, Yeongmin, et al.
Veröffentlicht: (2026)