Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jelassi, Samy, Kwun, Mujin, Zhao, Rosie, Li, Yuanzhi, Fusi, Nicolo, Du, Yilun, Kakade, Sham M., Domingo-Enrich, Carles
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912968215625728
author Jelassi, Samy
Kwun, Mujin
Zhao, Rosie
Li, Yuanzhi
Fusi, Nicolo
Du, Yilun
Kakade, Sham M.
Domingo-Enrich, Carles
author_facet Jelassi, Samy
Kwun, Mujin
Zhao, Rosie
Li, Yuanzhi
Fusi, Nicolo
Du, Yilun
Kakade, Sham M.
Domingo-Enrich, Carles
contents Cross-entropy (CE) training provides dense and scalable supervision for language models, but it optimizes next-token prediction under teacher forcing rather than sequence-level behavior under model rollouts. We introduce a feature-matching objective for language-model fine-tuning that targets sequence-level statistics of the completion distribution, providing dense semantic feedback without requiring a task-specific verifier or preference model. To optimize this objective efficiently, we propose energy-based fine-tuning (EBFT), which uses strided block-parallel sampling to generate multiple rollouts from nested prefixes concurrently, batches feature extraction over these rollouts, and uses the resulting embeddings to perform an on-policy policy-gradient update. We present a theoretical perspective connecting EBFT to KL-regularized feature-matching and energy-based modeling. Empirically, across Q&A coding, unstructured coding, and translation, EBFT matches RLVR and outperforms SFT on downstream accuracy while achieving a lower validation cross-entropy than both methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12248
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
Jelassi, Samy
Kwun, Mujin
Zhao, Rosie
Li, Yuanzhi
Fusi, Nicolo
Du, Yilun
Kakade, Sham M.
Domingo-Enrich, Carles
Machine Learning
Cross-entropy (CE) training provides dense and scalable supervision for language models, but it optimizes next-token prediction under teacher forcing rather than sequence-level behavior under model rollouts. We introduce a feature-matching objective for language-model fine-tuning that targets sequence-level statistics of the completion distribution, providing dense semantic feedback without requiring a task-specific verifier or preference model. To optimize this objective efficiently, we propose energy-based fine-tuning (EBFT), which uses strided block-parallel sampling to generate multiple rollouts from nested prefixes concurrently, batches feature extraction over these rollouts, and uses the resulting embeddings to perform an on-policy policy-gradient update. We present a theoretical perspective connecting EBFT to KL-regularized feature-matching and energy-based modeling. Empirically, across Q&A coding, unstructured coding, and translation, EBFT matches RLVR and outperforms SFT on downstream accuracy while achieving a lower validation cross-entropy than both methods.
title Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
topic Machine Learning
url https://arxiv.org/abs/2603.12248