Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Yuxin, Huang, Bo, Wang, Yufei, Zeng, Xingshan, Li, Liangyou, Wang, Yasheng, Jiang, Xin, Shang, Lifeng, Tang, Ruiming, Wang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
ARTIS: Agentic Risk-Aware Test-Time Scaling via Iterative Simulation
von: Zeng, Xingshan, et al.
Veröffentlicht: (2026)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2026)
Learning to Edit: Aligning LLMs with Knowledge Editing
von: Jiang, Yuxin, et al.
Veröffentlicht: (2024)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2024)
FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
von: Jiang, Yuxin, et al.
Veröffentlicht: (2023)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2023)
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
von: Wang, Zezhong, et al.
Veröffentlicht: (2025)
von: Wang, Zezhong, et al.
Veröffentlicht: (2025)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2024)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2024)
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2023)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2023)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
von: Huang, Xu, et al.
Veröffentlicht: (2025)
von: Huang, Xu, et al.
Veröffentlicht: (2025)
DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization
von: Deng, Mengyi, et al.
Veröffentlicht: (2026)
von: Deng, Mengyi, et al.
Veröffentlicht: (2026)
Position: The Real Barrier to LLM Agent Usability is Agentic ROI
von: Liu, Weiwen, et al.
Veröffentlicht: (2025)
von: Liu, Weiwen, et al.
Veröffentlicht: (2025)
ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution
von: Huang, Xu, et al.
Veröffentlicht: (2025)
von: Huang, Xu, et al.
Veröffentlicht: (2025)
Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
von: Huang, Shijue, et al.
Veröffentlicht: (2024)
von: Huang, Shijue, et al.
Veröffentlicht: (2024)
SELF: Self-Evolution with Language Feedback
von: Lu, Jianqiao, et al.
Veröffentlicht: (2023)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2023)
Data Management For Training Large Language Models: A Survey
von: Wang, Zige, et al.
Veröffentlicht: (2023)
von: Wang, Zige, et al.
Veröffentlicht: (2023)
Entropy Law: The Story Behind Data Compression and LLM Performance
von: Yin, Mingjia, et al.
Veröffentlicht: (2024)
von: Yin, Mingjia, et al.
Veröffentlicht: (2024)
YODA: Teacher-Student Progressive Learning for Language Models
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
NILE: Internal Consistency Alignment in Large Language Models
von: Hu, Minda, et al.
Veröffentlicht: (2024)
von: Hu, Minda, et al.
Veröffentlicht: (2024)
Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
von: Xu, Kaishuai, et al.
Veröffentlicht: (2024)
von: Xu, Kaishuai, et al.
Veröffentlicht: (2024)
Evaluating the External and Parametric Knowledge Fusion of Large Language Models
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
von: Gao, Fan, et al.
Veröffentlicht: (2025)
von: Gao, Fan, et al.
Veröffentlicht: (2025)
LogitsCoder: Towards Efficient Chain-of-Thought Path Search via Logits Preference Decoding for Code Generation
von: Chen, Jizheng, et al.
Veröffentlicht: (2026)
von: Chen, Jizheng, et al.
Veröffentlicht: (2026)
A Survey on Multi-Turn Interaction Capabilities of Large Language Models
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
von: Du, Yiming, et al.
Veröffentlicht: (2025)
von: Du, Yiming, et al.
Veröffentlicht: (2025)
Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
von: Jia, Xinda, et al.
Veröffentlicht: (2025)
von: Jia, Xinda, et al.
Veröffentlicht: (2025)
PrefPO: Pairwise Preference Prompt Optimization
von: Singhal, Rahul, et al.
Veröffentlicht: (2026)
von: Singhal, Rahul, et al.
Veröffentlicht: (2026)
Mitigating Large Language Model Hallucination with Faithful Finetuning
von: Hu, Minda, et al.
Veröffentlicht: (2024)
von: Hu, Minda, et al.
Veröffentlicht: (2024)
Learning to Align Multi-Faceted Evaluation: A Unified and Robust Framework
von: Xu, Kaishuai, et al.
Veröffentlicht: (2025)
von: Xu, Kaishuai, et al.
Veröffentlicht: (2025)
Token-level Direct Preference Optimization
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2024)
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2024)
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
von: Zhao, Wenshuo, et al.
Veröffentlicht: (2026)
von: Zhao, Wenshuo, et al.
Veröffentlicht: (2026)
ToolACE: Winning the Points of LLM Function Calling
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
von: Li, Xiangyang, et al.
Veröffentlicht: (2024)
von: Li, Xiangyang, et al.
Veröffentlicht: (2024)
Direct Judgement Preference Optimization
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025) -
ARTIS: Agentic Risk-Aware Test-Time Scaling via Iterative Simulation
von: Zeng, Xingshan, et al.
Veröffentlicht: (2026) -
Learning to Edit: Aligning LLMs with Knowledge Editing
von: Jiang, Yuxin, et al.
Veröffentlicht: (2024) -
FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
von: Jiang, Yuxin, et al.
Veröffentlicht: (2023) -
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)