Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Sen, Cui, Leyang, Cai, Deng, Huang, Xinting, Shi, Shuming, Lam, Wai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
by: Shi, Shuming, et al.
Published: (2024)
by: Shi, Shuming, et al.
Published: (2024)
Knowledge Verification to Nip Hallucination in the Bud
by: Wan, Fanqi, et al.
Published: (2024)
by: Wan, Fanqi, et al.
Published: (2024)
Neuro-Symbolic Integration Brings Causal and Reliable Reasoning Proofs
by: Yang, Sen, et al.
Published: (2023)
by: Yang, Sen, et al.
Published: (2023)
Reasons to Reject? Aligning Language Models with Judgments
by: Xu, Weiwen, et al.
Published: (2023)
by: Xu, Weiwen, et al.
Published: (2023)
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving
by: Gao, Chang, et al.
Published: (2023)
by: Gao, Chang, et al.
Published: (2023)
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
by: Chung, Tsz Ting, et al.
Published: (2024)
by: Chung, Tsz Ting, et al.
Published: (2024)
A Frustratingly Simple Decoding Method for Neural Text Generation
by: Yang, Haoran, et al.
Published: (2023)
by: Yang, Haoran, et al.
Published: (2023)
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
Knowledge Fusion of Large Language Models
by: Wan, Fanqi, et al.
Published: (2024)
by: Wan, Fanqi, et al.
Published: (2024)
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
by: Zhang, Yue, et al.
Published: (2023)
by: Zhang, Yue, et al.
Published: (2023)
Retrieval is Accurate Generation
by: Cao, Bowen, et al.
Published: (2024)
by: Cao, Bowen, et al.
Published: (2024)
LRHP: Learning Representations for Human Preferences via Preference Pairs
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning
by: Zhang, Dongxu, et al.
Published: (2025)
by: Zhang, Dongxu, et al.
Published: (2025)
Users as Annotators: LLM Preference Learning from Comparison Mode
by: Cai, Zhongze, et al.
Published: (2025)
by: Cai, Zhongze, et al.
Published: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction Wrapping
by: Chen, Yongrui, et al.
Published: (2023)
by: Chen, Yongrui, et al.
Published: (2023)
Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
by: Wang, Yao, et al.
Published: (2025)
by: Wang, Yao, et al.
Published: (2025)
On the Transformations across Reward Model, Parameter Update, and In-Context Prompt
by: Cai, Deng, et al.
Published: (2024)
by: Cai, Deng, et al.
Published: (2024)
Iterative Reasoning Preference Optimization
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
by: Li, Yafu, et al.
Published: (2024)
by: Li, Yafu, et al.
Published: (2024)
Tokenization Preference for Human and Machine Learning Model: An Annotation Study
by: Hiraoka, Tatsuya, et al.
Published: (2023)
by: Hiraoka, Tatsuya, et al.
Published: (2023)
No Preference Left Behind: Group Distributional Preference Optimization
by: Yao, Binwei, et al.
Published: (2024)
by: Yao, Binwei, et al.
Published: (2024)
All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing
by: Deshmukh, Advait, et al.
Published: (2024)
by: Deshmukh, Advait, et al.
Published: (2024)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
by: Agrawal, Sweta, et al.
Published: (2024)
by: Agrawal, Sweta, et al.
Published: (2024)
Enhancing the Preference Extractor in Multi-turn Dialogues: From Annotating Disasters to Accurate Preference Extraction
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Not All Options Are Created Equal: Textual Option Weighting for Token-Efficient LLM-Based Knowledge Tracing
by: Kim, JongWoo, et al.
Published: (2024)
by: Kim, JongWoo, et al.
Published: (2024)
InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformation
by: Cao, Bowen, et al.
Published: (2025)
by: Cao, Bowen, et al.
Published: (2025)
PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs
by: Zhang, Rongzhi, et al.
Published: (2024)
by: Zhang, Rongzhi, et al.
Published: (2024)
AIPO: Improving Training Objective for Iterative Preference Optimization
by: Shen, Yaojie, et al.
Published: (2024)
by: Shen, Yaojie, et al.
Published: (2024)
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
by: Zhang, Yue, et al.
Published: (2023)
by: Zhang, Yue, et al.
Published: (2023)
DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
by: Huang, Chengyu, et al.
Published: (2025)
by: Huang, Chengyu, et al.
Published: (2025)
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
by: Feng, Duanyu, et al.
Published: (2024)
by: Feng, Duanyu, et al.
Published: (2024)
Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
by: Huang, Jianheng, et al.
Published: (2024)
by: Huang, Jianheng, et al.
Published: (2024)
Not All Preferences Are Created Equal: Stability-Aware and Gradient-Efficient Alignment for Reasoning Models
by: Wu, Hui, et al.
Published: (2026)
by: Wu, Hui, et al.
Published: (2026)
A Thorough Examination of Decoding Methods in the Era of LLMs
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
by: Huang, Kun, et al.
Published: (2025)
by: Huang, Kun, et al.
Published: (2025)
MAGE: Machine-generated Text Detection in the Wild
by: Li, Yafu, et al.
Published: (2023)
by: Li, Yafu, et al.
Published: (2023)
Preference-Aware Rubric Learning for Personalized Evaluation
by: Qiu, Yilun, et al.
Published: (2026)
by: Qiu, Yilun, et al.
Published: (2026)
DepWiGNN: A Depth-wise Graph Neural Network for Multi-hop Spatial Reasoning in Text
by: Li, Shuaiyi, et al.
Published: (2023)
by: Li, Shuaiyi, et al.
Published: (2023)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Similar Items
-
Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
by: Shi, Shuming, et al.
Published: (2024) -
Knowledge Verification to Nip Hallucination in the Bud
by: Wan, Fanqi, et al.
Published: (2024) -
Neuro-Symbolic Integration Brings Causal and Reliable Reasoning Proofs
by: Yang, Sen, et al.
Published: (2023) -
Reasons to Reject? Aligning Language Models with Judgments
by: Xu, Weiwen, et al.
Published: (2023) -
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving
by: Gao, Chang, et al.
Published: (2023)