SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917539319119872 |
|---|---|
| author | Yoon, Kanghoon Kim, Minsub Lee, Sungjae Lee, Joonhyung Woo, Sunghyeon In, Yeonjun Kwon, Se Jung Park, Chanyoung Lee, Dongsoo |
| author_facet | Yoon, Kanghoon Kim, Minsub Lee, Sungjae Lee, Joonhyung Woo, Sunghyeon In, Yeonjun Kwon, Se Jung Park, Chanyoung Lee, Dongsoo |
| contents | Speculative decoding accelerates LLM inference by verifying candidate tokens from a draft model against a larger target model. Recent judge decoding boosts this process by relaxing verification criteria by accepting draft tokens that may exhibit minor discrepancies from target model output, but existing methods are restricted by their reliance on human annotations or tasks with verifiable ground truths, limiting generalizability across diverse NLP tasks. We propose SelfJudge, which trains judge verifiers via self-supervision of the target model. Our method measures semantic preservation by assessing whether token-substituted responses preserve the meaning of original responses, enabling automatic verifier training across diverse NLP tasks. Our experiments show SelfJudge achieves superior inference-accuracy trade-offs than judge decoding baselines, offering a broadly applicable solution for faster LLM inference. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_02329 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification Yoon, Kanghoon Kim, Minsub Lee, Sungjae Lee, Joonhyung Woo, Sunghyeon In, Yeonjun Kwon, Se Jung Park, Chanyoung Lee, Dongsoo Computation and Language Artificial Intelligence Speculative decoding accelerates LLM inference by verifying candidate tokens from a draft model against a larger target model. Recent judge decoding boosts this process by relaxing verification criteria by accepting draft tokens that may exhibit minor discrepancies from target model output, but existing methods are restricted by their reliance on human annotations or tasks with verifiable ground truths, limiting generalizability across diverse NLP tasks. We propose SelfJudge, which trains judge verifiers via self-supervision of the target model. Our method measures semantic preservation by assessing whether token-substituted responses preserve the meaning of original responses, enabling automatic verifier training across diverse NLP tasks. Our experiments show SelfJudge achieves superior inference-accuracy trade-offs than judge decoding baselines, offering a broadly applicable solution for faster LLM inference. |
| title | SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.02329 |