TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Jiahao, Lu, Yifu, Zeng, Yifan, Guo, Jiacheng, Geng, Jiayi, Zhu, Chenhao, Juan, Xinzhe, Yang, Ling, Wang, Huazheng, Huang, Kaixuan, Wu, Yue, Wang, Mengdi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918134571597824
author Qiu, Jiahao
Lu, Yifu
Zeng, Yifan
Guo, Jiacheng
Geng, Jiayi
Zhu, Chenhao
Juan, Xinzhe
Yang, Ling
Wang, Huazheng
Huang, Kaixuan
Wu, Yue
Wang, Mengdi
author_facet Qiu, Jiahao
Lu, Yifu
Zeng, Yifan
Guo, Jiacheng
Geng, Jiayi
Zhu, Chenhao
Juan, Xinzhe
Yang, Ling
Wang, Huazheng
Huang, Kaixuan
Wu, Yue
Wang, Mengdi
contents Inference-time alignment enhances the performance of large language models without requiring additional training or fine-tuning but presents challenges due to balancing computational efficiency with high-quality output. Best-of-N (BoN) sampling, as a simple yet powerful approach, generates multiple responses and selects the best one, achieving improved performance but with a high computational cost. We propose TreeBoN, a novel framework that integrates a speculative tree-search strategy into Best-of-N (BoN) Sampling. TreeBoN maintains a set of parent nodes, iteratively branching and pruning low-quality responses, thereby reducing computational overhead while maintaining high output quality. Our approach also leverages token-level rewards from Direct Preference Optimization (DPO) to guide tree expansion and prune low-quality paths. We evaluate TreeBoN using AlpacaFarm, HH-RLHF, UltraFeedback, GSM8K, and TutorEval datasets, demonstrating consistent improvements. Specifically, TreeBoN achieves the highest win rate of 65% on TutorEval and around 60% win rates across other different datasets, outperforming standard BoN with the same computational cost and showcasing its scalability and alignment efficacy.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16033
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
Qiu, Jiahao
Lu, Yifu
Zeng, Yifan
Guo, Jiacheng
Geng, Jiayi
Zhu, Chenhao
Juan, Xinzhe
Yang, Ling
Wang, Huazheng
Huang, Kaixuan
Wu, Yue
Wang, Mengdi
Computation and Language
Artificial Intelligence
Machine Learning
Inference-time alignment enhances the performance of large language models without requiring additional training or fine-tuning but presents challenges due to balancing computational efficiency with high-quality output. Best-of-N (BoN) sampling, as a simple yet powerful approach, generates multiple responses and selects the best one, achieving improved performance but with a high computational cost. We propose TreeBoN, a novel framework that integrates a speculative tree-search strategy into Best-of-N (BoN) Sampling. TreeBoN maintains a set of parent nodes, iteratively branching and pruning low-quality responses, thereby reducing computational overhead while maintaining high output quality. Our approach also leverages token-level rewards from Direct Preference Optimization (DPO) to guide tree expansion and prune low-quality paths. We evaluate TreeBoN using AlpacaFarm, HH-RLHF, UltraFeedback, GSM8K, and TutorEval datasets, demonstrating consistent improvements. Specifically, TreeBoN achieves the highest win rate of 65% on TutorEval and around 60% win rates across other different datasets, outperforming standard BoN with the same computational cost and showcasing its scalability and alignment efficacy.
title TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.16033