Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Zhang, Zhendong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917917406265344
author Zhang, Zhendong
author_facet Zhang, Zhendong
contents Multiple heads decoding accelerates the inference of Large Language Models (LLMs) by predicting next several tokens simultaneously. It generates and verifies multiple candidate sequences in parallel via tree attention with a fixed structure. In this paper, we replace the fixed tree attention with dynamic tree attention on multiple head decoding, specifically in the context of MEDUSA. We propose a simple and low complexity strategy to generate candidates and construct the dynamic tree structure. Preliminary experiments show that the proposed method improves the decoding efficiency of multiple head decoding for LLMs while maintaining the generation quality. This result demonstrates the potential for improvement of multiple head decoding in candidate generation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05947
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
Zhang, Zhendong
Computer Vision and Pattern Recognition
Computation and Language
Multiple heads decoding accelerates the inference of Large Language Models (LLMs) by predicting next several tokens simultaneously. It generates and verifies multiple candidate sequences in parallel via tree attention with a fixed structure. In this paper, we replace the fixed tree attention with dynamic tree attention on multiple head decoding, specifically in the context of MEDUSA. We propose a simple and low complexity strategy to generate candidates and construct the dynamic tree structure. Preliminary experiments show that the proposed method improves the decoding efficiency of multiple head decoding for LLMs while maintaining the generation quality. This result demonstrates the potential for improvement of multiple head decoding in candidate generation.
title Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2502.05947