BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Sizhe, Tong, Yongqi, Zhang, Hengyuan, Li, Dawei, Zhang, Xin, Chen, Tianlong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BPO: Revisiting Preference Modeling in Direct Preference Optimization
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning
von: Tong, Yongqi, et al.
Veröffentlicht: (2024)
von: Tong, Yongqi, et al.
Veröffentlicht: (2024)
Optimizing Language Model's Reasoning Abilities with Weak Supervision
von: Tong, Yongqi, et al.
Veröffentlicht: (2024)
von: Tong, Yongqi, et al.
Veröffentlicht: (2024)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
von: Robertson, Alex, et al.
Veröffentlicht: (2026)
von: Robertson, Alex, et al.
Veröffentlicht: (2026)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
REWARD CONSISTENCY: Improving Multi-Objective Alignment from a Data-Centric Perspective
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
Contextualization Distillation from Large Language Model for Knowledge Graph Completion
von: Li, Dawei, et al.
Veröffentlicht: (2024)
von: Li, Dawei, et al.
Veröffentlicht: (2024)
Establishing Knowledge Preference in Language Models
von: Zhou, Sizhe, et al.
Veröffentlicht: (2024)
von: Zhou, Sizhe, et al.
Veröffentlicht: (2024)
Balancing Speciality and Versatility: A Coarse to Fine Framework for Mitigating Catastrophic Forgetting in Large Language Models
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2024)
Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering
von: Zhang, Yichi, et al.
Veröffentlicht: (2023)
von: Zhang, Yichi, et al.
Veröffentlicht: (2023)
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs
von: Ding, Yue, et al.
Veröffentlicht: (2025)
von: Ding, Yue, et al.
Veröffentlicht: (2025)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
von: Chen, Xuetian, et al.
Veröffentlicht: (2025)
von: Chen, Xuetian, et al.
Veröffentlicht: (2025)
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization
von: Song, Feifan, et al.
Veröffentlicht: (2024)
von: Song, Feifan, et al.
Veröffentlicht: (2024)
MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization
von: Lyu, Yougang, et al.
Veröffentlicht: (2024)
von: Lyu, Yougang, et al.
Veröffentlicht: (2024)
Preference Ranking Optimization for Human Alignment
von: Song, Feifan, et al.
Veröffentlicht: (2023)
von: Song, Feifan, et al.
Veröffentlicht: (2023)
Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
von: Li, Jian, et al.
Veröffentlicht: (2025)
von: Li, Jian, et al.
Veröffentlicht: (2025)
Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
von: Li, Tong, et al.
Veröffentlicht: (2025)
von: Li, Tong, et al.
Veröffentlicht: (2025)
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization
von: Wang, Yunan, et al.
Veröffentlicht: (2025)
von: Wang, Yunan, et al.
Veröffentlicht: (2025)
StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models
von: Khan, Ishmam, et al.
Veröffentlicht: (2026)
von: Khan, Ishmam, et al.
Veröffentlicht: (2026)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
von: Guo, Yiju, et al.
Veröffentlicht: (2024)
von: Guo, Yiju, et al.
Veröffentlicht: (2024)
Beyond Entity Alignment: Towards Complete Knowledge Graph Alignment via Entity-Relation Synergy
von: Fang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Fang, Xiaohan, et al.
Veröffentlicht: (2024)
Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
von: Cui, Chaoqun, et al.
Veröffentlicht: (2025)
von: Cui, Chaoqun, et al.
Veröffentlicht: (2025)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
Dual Engines of Thoughts: A Depth-Breadth Integration Framework for Open-Ended Analysis
von: Yu, Fei-Hsuan, et al.
Veröffentlicht: (2025)
von: Yu, Fei-Hsuan, et al.
Veröffentlicht: (2025)
Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
von: Wang, Mingzhi, et al.
Veröffentlicht: (2024)
von: Wang, Mingzhi, et al.
Veröffentlicht: (2024)
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Knowledge Graph Construction in Power Distribution Networks
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2024)
Dynamic Fault Analysis in Substations Based on Knowledge Graphs
von: Li, Weiwei, et al.
Veröffentlicht: (2023)
von: Li, Weiwei, et al.
Veröffentlicht: (2023)
Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge
von: Sui, Yi, et al.
Veröffentlicht: (2025)
von: Sui, Yi, et al.
Veröffentlicht: (2025)
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
Dual Reasoning: A GNN-LLM Collaborative Framework for Knowledge Graph Question Answering
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
von: Li, Moxin, et al.
Veröffentlicht: (2025)
von: Li, Moxin, et al.
Veröffentlicht: (2025)
Panacea: Pareto Alignment via Preference Adaptation for LLMs
von: Zhong, Yifan, et al.
Veröffentlicht: (2024)
von: Zhong, Yifan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BPO: Revisiting Preference Modeling in Direct Preference Optimization
von: Sun, Lin, et al.
Veröffentlicht: (2025) -
Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning
von: Tong, Yongqi, et al.
Veröffentlicht: (2024) -
Optimizing Language Model's Reasoning Abilities with Weak Supervision
von: Tong, Yongqi, et al.
Veröffentlicht: (2024) -
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
von: Xu, Wenda, et al.
Veröffentlicht: (2024) -
KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
von: Robertson, Alex, et al.
Veröffentlicht: (2026)