Advancing LLM Reasoning Generalists with Preference Trees
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yuan, Lifan, Cui, Ganqu, Wang, Hanbin, Ding, Ning, Wang, Xingyao, Deng, Jia, Shan, Boji, Chen, Huimin, Xie, Ruobing, Lin, Yankai, Liu, Zhenghao, Zhou, Bowen, Peng, Hao, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
par: Guo, Yiju, et autres
Publié: (2024)
par: Guo, Yiju, et autres
Publié: (2024)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
par: Wang, Hanbin, et autres
Publié: (2023)
par: Wang, Hanbin, et autres
Publié: (2023)
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
par: Yuan, Lifan, et autres
Publié: (2025)
par: Yuan, Lifan, et autres
Publié: (2025)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
par: Cui, Ganqu, et autres
Publié: (2023)
par: Cui, Ganqu, et autres
Publié: (2023)
Mastering Text, Code and Math Simultaneously via Fusing Highly Specialized Language Models
par: Ding, Ning, et autres
Publié: (2024)
par: Ding, Ning, et autres
Publié: (2024)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
par: He, Bingxiang, et autres
Publié: (2024)
par: He, Bingxiang, et autres
Publié: (2024)
Free Process Rewards without Process Labels
par: Yuan, Lifan, et autres
Publié: (2024)
par: Yuan, Lifan, et autres
Publié: (2024)
UltraMedical: Building Specialized Generalists in Biomedicine
par: Zhang, Kaiyan, et autres
Publié: (2024)
par: Zhang, Kaiyan, et autres
Publié: (2024)
Representation Learning for Natural Language Processing
par: Liu, Zhiyuan, et autres
Publié: (2020)
par: Liu, Zhiyuan, et autres
Publié: (2020)
RLPR: Extrapolating RLVR to General Domains without Verifiers
par: Yu, Tianyu, et autres
Publié: (2025)
par: Yu, Tianyu, et autres
Publié: (2025)
Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
par: Gao, Cheng, et autres
Publié: (2024)
par: Gao, Cheng, et autres
Publié: (2024)
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
par: He, Bingxiang, et autres
Publié: (2025)
par: He, Bingxiang, et autres
Publié: (2025)
PersLLM: A Personified Training Approach for Large Language Models
par: Zeng, Zheni, et autres
Publié: (2024)
par: Zeng, Zheni, et autres
Publié: (2024)
Teaching Large Reasoning Models Effective Reflection
par: Wang, Hanbin, et autres
Publié: (2026)
par: Wang, Hanbin, et autres
Publié: (2026)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
par: Cui, Ganqu, et autres
Publié: (2025)
par: Cui, Ganqu, et autres
Publié: (2025)
KBAlign: Efficient Self Adaptation on Specific Knowledge Bases
par: Zeng, Zheni, et autres
Publié: (2024)
par: Zeng, Zheni, et autres
Publié: (2024)
KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs
par: Wu, Dingjun, et autres
Publié: (2025)
par: Wu, Dingjun, et autres
Publié: (2025)
The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
par: Chen, Weize, et autres
Publié: (2025)
par: Chen, Weize, et autres
Publié: (2025)
Empowering Private Tutoring by Chaining Large Language Models
par: Chen, Yulin, et autres
Publié: (2023)
par: Chen, Yulin, et autres
Publié: (2023)
Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
par: Liu, Genglin, et autres
Publié: (2023)
par: Liu, Genglin, et autres
Publié: (2023)
Exploring the Benefit of Activation Sparsity in Pre-training
par: Zhang, Zhengyan, et autres
Publié: (2024)
par: Zhang, Zhengyan, et autres
Publié: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
par: Lv, Xingtai, et autres
Publié: (2024)
par: Lv, Xingtai, et autres
Publié: (2024)
A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
par: Xu, Xiaoang, et autres
Publié: (2025)
par: Xu, Xiaoang, et autres
Publié: (2025)
UltraIF: Advancing Instruction Following from the Wild
par: An, Kaikai, et autres
Publié: (2025)
par: An, Kaikai, et autres
Publié: (2025)
KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG
par: Li, Yongjian, et autres
Publié: (2025)
par: Li, Yongjian, et autres
Publié: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
par: Sun, Yubo, et autres
Publié: (2025)
par: Sun, Yubo, et autres
Publié: (2025)
DeepNote: Note-Centric Deep Retrieval-Augmented Generation
par: Wang, Ruobing, et autres
Publié: (2024)
par: Wang, Ruobing, et autres
Publié: (2024)
NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
par: Chen, Huayu, et autres
Publié: (2025)
par: Chen, Huayu, et autres
Publié: (2025)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
par: Duan, Shaohua, et autres
Publié: (2025)
par: Duan, Shaohua, et autres
Publié: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
par: Xiao, Chaojun, et autres
Publié: (2024)
par: Xiao, Chaojun, et autres
Publié: (2024)
Process Reinforcement through Implicit Rewards
par: Cui, Ganqu, et autres
Publié: (2025)
par: Cui, Ganqu, et autres
Publié: (2025)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
par: Soni, Aditya Bharat, et autres
Publié: (2025)
par: Soni, Aditya Bharat, et autres
Publié: (2025)
A representation of range decreasing group homomorphisms
par: Zhang, Ning, et autres
Publié: (2025)
par: Zhang, Ning, et autres
Publié: (2025)
Executable Code Actions Elicit Better LLM Agents
par: Wang, Xingyao, et autres
Publié: (2024)
par: Wang, Xingyao, et autres
Publié: (2024)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
par: Cheng, Qianjia, et autres
Publié: (2026)
par: Cheng, Qianjia, et autres
Publié: (2026)
Noise Contrastive Alignment of Language Models with Explicit Rewards
par: Chen, Huayu, et autres
Publié: (2024)
par: Chen, Huayu, et autres
Publié: (2024)
UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
par: Wang, Haoyu, et autres
Publié: (2024)
par: Wang, Haoyu, et autres
Publié: (2024)
Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication
par: Chen, Weize, et autres
Publié: (2024)
par: Chen, Weize, et autres
Publié: (2024)
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
par: He, Bingxiang, et autres
Publié: (2025)
par: He, Bingxiang, et autres
Publié: (2025)
WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
par: Fan, Shengda, et autres
Publié: (2024)
par: Fan, Shengda, et autres
Publié: (2024)
Documents similaires
-
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
par: Guo, Yiju, et autres
Publié: (2024) -
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
par: Wang, Hanbin, et autres
Publié: (2023) -
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
par: Yuan, Lifan, et autres
Publié: (2025) -
UltraFeedback: Boosting Language Models with Scaled AI Feedback
par: Cui, Ganqu, et autres
Publié: (2023) -
Mastering Text, Code and Math Simultaneously via Fusing Highly Specialized Language Models
par: Ding, Ning, et autres
Publié: (2024)