Towards Understanding Valuable Preference Data for Large Language Model Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zizhuo, Wang, Qizhou, Ye, Shanshan, Zhu, Jianing, Yao, Jiangchao, Han, Bo, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Is Preference Optimization Doing, and Why?
von: Wang, Yue, et al.
Veröffentlicht: (2025)
von: Wang, Yue, et al.
Veröffentlicht: (2025)
Towards Effective Evaluations and Comparisons for LLM Unlearning Methods
von: Wang, Qizhou, et al.
Veröffentlicht: (2024)
von: Wang, Qizhou, et al.
Veröffentlicht: (2024)
Decoupling the Class Label and the Target Concept in Machine Unlearning
von: Zhu, Jianing, et al.
Veröffentlicht: (2024)
von: Zhu, Jianing, et al.
Veröffentlicht: (2024)
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
von: Zhang, Zizhuo, et al.
Veröffentlicht: (2025)
von: Zhang, Zizhuo, et al.
Veröffentlicht: (2025)
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
von: Li, Xuan, et al.
Veröffentlicht: (2023)
von: Li, Xuan, et al.
Veröffentlicht: (2023)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
von: Yu, Geng, et al.
Veröffentlicht: (2024)
von: Yu, Geng, et al.
Veröffentlicht: (2024)
Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
von: Taheri, Ali, et al.
Veröffentlicht: (2025)
von: Taheri, Ali, et al.
Veröffentlicht: (2025)
Fast and Accurate Blind Flexible Docking
von: Zhang, Zizhuo, et al.
Veröffentlicht: (2025)
von: Zhang, Zizhuo, et al.
Veröffentlicht: (2025)
Towards Understanding How Knowledge Evolves in Large Vision-Language Models
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
Per-parameter Task Arithmetic for Unlearning in Large Language Models
von: Cai, Chengyi, et al.
Veröffentlicht: (2026)
von: Cai, Chengyi, et al.
Veröffentlicht: (2026)
Towards Scalable Oversight with Collaborative Multi-Agent Debate in Error Detection
von: Chen, Yongqiang, et al.
Veröffentlicht: (2025)
von: Chen, Yongqiang, et al.
Veröffentlicht: (2025)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
Is Gradient Ascent Really Necessary? Memorize to Forget for Machine Unlearning
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware Minimization
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
From Coefficients to Directions: Rethinking Model Merging with Directional Alignment
von: Chen, Zhikang, et al.
Veröffentlicht: (2025)
von: Chen, Zhikang, et al.
Veröffentlicht: (2025)
Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
von: Gu, Zhuojun, et al.
Veröffentlicht: (2025)
von: Gu, Zhuojun, et al.
Veröffentlicht: (2025)
Less is More: One-shot Subgraph Reasoning on Large-scale Knowledge Graphs
von: Zhou, Zhanke, et al.
Veröffentlicht: (2024)
von: Zhou, Zhanke, et al.
Veröffentlicht: (2024)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
von: Yang, Puning, et al.
Veröffentlicht: (2026)
von: Yang, Puning, et al.
Veröffentlicht: (2026)
Federated Learning with Bilateral Curation for Partially Class-Disjoint Data
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
FedPDPO: Federated Personalized Direct Preference Optimization for Large Language Model Alignment
von: Zhu, Kewen, et al.
Veröffentlicht: (2026)
von: Zhu, Kewen, et al.
Veröffentlicht: (2026)
A Fast Algorithm for the Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit
von: Nakamura, Shintaro, et al.
Veröffentlicht: (2023)
von: Nakamura, Shintaro, et al.
Veröffentlicht: (2023)
Enriching Disentanglement: From Logical Definitions to Quantitative Metrics
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
Differentially Private Preference Data Synthesis for Large Language Model Alignment
von: Gao, Fengyu, et al.
Veröffentlicht: (2026)
von: Gao, Fengyu, et al.
Veröffentlicht: (2026)
Towards Scalable Oversight via Partitioned Human Supervision
von: Yin, Ren, et al.
Veröffentlicht: (2025)
von: Yin, Ren, et al.
Veröffentlicht: (2025)
Multi-Player Approaches for Dueling Bandits
von: Raveh, Or, et al.
Veröffentlicht: (2024)
von: Raveh, Or, et al.
Veröffentlicht: (2024)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
von: Wen, Rui, et al.
Veröffentlicht: (2024)
von: Wen, Rui, et al.
Veröffentlicht: (2024)
OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
A Category-theoretical Meta-analysis of Definitions of Disentanglement
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
von: Zhang, Yivan, et al.
Veröffentlicht: (2023)
Unraveling the Impact of Heterophilic Structures on Graph Positive-Unlabeled Learning
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
Tangent Space Fine-Tuning for Directional Preference Alignment in Large Language Models
von: Erdogan, Mete
Veröffentlicht: (2026)
von: Erdogan, Mete
Veröffentlicht: (2026)
Accelerated Preference Optimization for Large Language Model Alignment
von: He, Jiafan, et al.
Veröffentlicht: (2024)
von: He, Jiafan, et al.
Veröffentlicht: (2024)
Data Selection for LLM Alignment Using Fine-Grained Preferences
von: Zhang, Jia, et al.
Veröffentlicht: (2025)
von: Zhang, Jia, et al.
Veröffentlicht: (2025)
VEC-SBM: Optimal Community Detection with Vectorial Edges Covariates
von: Braun, Guillaume, et al.
Veröffentlicht: (2024)
von: Braun, Guillaume, et al.
Veröffentlicht: (2024)
Riemannian Langevin Dynamics: Strong Convergence of Geometric Euler-Maruyama Scheme
von: Zhan, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhan, Zhiyuan, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning with Domain-Unlabeled Data
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2024)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2024)
Doubly Robust Alignment for Large Language Models
von: Xu, Erhan, et al.
Veröffentlicht: (2025)
von: Xu, Erhan, et al.
Veröffentlicht: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
von: Li, Fengpeng, et al.
Veröffentlicht: (2026)
von: Li, Fengpeng, et al.
Veröffentlicht: (2026)
Can Large Language Models Understand Intermediate Representations in Compilers?
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What Is Preference Optimization Doing, and Why?
von: Wang, Yue, et al.
Veröffentlicht: (2025) -
Towards Effective Evaluations and Comparisons for LLM Unlearning Methods
von: Wang, Qizhou, et al.
Veröffentlicht: (2024) -
Decoupling the Class Label and the Target Concept in Machine Unlearning
von: Zhu, Jianing, et al.
Veröffentlicht: (2024) -
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
von: Zhang, Zizhuo, et al.
Veröffentlicht: (2025) -
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
von: Li, Xuan, et al.
Veröffentlicht: (2023)