Self-Speculative Biased Decoding for Faster Re-Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Linxiao, Deng, Haoyun, Shu, Kangyuan, Wang, Shizhen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908745859072000
author Zeng, Linxiao
Deng, Haoyun
Shu, Kangyuan
Wang, Shizhen
author_facet Zeng, Linxiao
Deng, Haoyun
Shu, Kangyuan
Wang, Shizhen
contents Large language models achieve strong machine translation quality but incur high inference cost and latency, posing challenges for simultaneous translation. Re-translation provides a practical solution for off-the-shelf LLMs by repeatedly regenerating the target output as the source input grows, but it suffers from substantial redundant computation. We propose Self-Speculative Biased Decoding (SSBD), a simple and tuning-free inference method that accelerates re-translation by exploiting temporal coherence in streaming translation. SSBD reuses the model's previous output as a speculative draft for the updated input, verifies the draft efficiently in a single forward pass with a lightweight bias, and resumes autoregressive decoding only from the first divergence. We further introduce a display-only masking strategy that hides unstable suffixes from the user interface while retaining them in the draft for verification and potential acceptance. Experiments show that SSBD achieves substantial speedup over standard re-translation while maintaining comparable translation quality, without architectural changes, auxiliary models, or extra fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Speculative Biased Decoding for Faster Re-Translation
Zeng, Linxiao
Deng, Haoyun
Shu, Kangyuan
Wang, Shizhen
Computation and Language
Artificial Intelligence
Machine Learning
Large language models achieve strong machine translation quality but incur high inference cost and latency, posing challenges for simultaneous translation. Re-translation provides a practical solution for off-the-shelf LLMs by repeatedly regenerating the target output as the source input grows, but it suffers from substantial redundant computation. We propose Self-Speculative Biased Decoding (SSBD), a simple and tuning-free inference method that accelerates re-translation by exploiting temporal coherence in streaming translation. SSBD reuses the model's previous output as a speculative draft for the updated input, verifies the draft efficiently in a single forward pass with a lightweight bias, and resumes autoregressive decoding only from the first divergence. We further introduce a display-only masking strategy that hides unstable suffixes from the user interface while retaining them in the draft for verification and potential acceptance. Experiments show that SSBD achieves substantial speedup over standard re-translation while maintaining comparable translation quality, without architectural changes, auxiliary models, or extra fine-tuning.
title Self-Speculative Biased Decoding for Faster Re-Translation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.21740