SLOT: Sample-specific Language Model Optimization at Test-time

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hu, Yang, Zhang, Xingyu, Fang, Xueji, Chen, Zhiyang, Wang, Xiao, Zhang, Huatian, Qi, Guojun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912394441129984
author Hu, Yang
Zhang, Xingyu
Fang, Xueji
Chen, Zhiyang
Wang, Xiao
Zhang, Huatian
Qi, Guojun
author_facet Hu, Yang
Zhang, Xingyu
Fang, Xueji
Chen, Zhiyang
Wang, Xiao
Zhang, Huatian
Qi, Guojun
contents We propose SLOT (Sample-specific Language Model Optimization at Test-time), a novel and parameter-efficient test-time inference approach that enhances a language model's ability to more accurately respond to individual prompts. Existing Large Language Models (LLMs) often struggle with complex instructions, leading to poor performances on those not well represented among general samples. To address this, SLOT conducts few optimization steps at test-time to update a light-weight sample-specific parameter vector. It is added to the final hidden layer before the output head, and enables efficient adaptation by caching the last layer features during per-sample optimization. By minimizing the cross-entropy loss on the input prompt only, SLOT helps the model better aligned with and follow each given instruction. In experiments, we demonstrate that our method outperforms the compared models across multiple benchmarks and LLMs. For example, Qwen2.5-7B with SLOT achieves an accuracy gain of 8.6% on GSM8K from 57.54% to 66.19%, while DeepSeek-R1-Distill-Llama-70B with SLOT achieves a SOTA accuracy of 68.69% on GPQA among 70B-level models. Our code is available at https://github.com/maple-research-lab/SLOT.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SLOT: Sample-specific Language Model Optimization at Test-time
Hu, Yang
Zhang, Xingyu
Fang, Xueji
Chen, Zhiyang
Wang, Xiao
Zhang, Huatian
Qi, Guojun
Computation and Language
Artificial Intelligence
Machine Learning
We propose SLOT (Sample-specific Language Model Optimization at Test-time), a novel and parameter-efficient test-time inference approach that enhances a language model's ability to more accurately respond to individual prompts. Existing Large Language Models (LLMs) often struggle with complex instructions, leading to poor performances on those not well represented among general samples. To address this, SLOT conducts few optimization steps at test-time to update a light-weight sample-specific parameter vector. It is added to the final hidden layer before the output head, and enables efficient adaptation by caching the last layer features during per-sample optimization. By minimizing the cross-entropy loss on the input prompt only, SLOT helps the model better aligned with and follow each given instruction. In experiments, we demonstrate that our method outperforms the compared models across multiple benchmarks and LLMs. For example, Qwen2.5-7B with SLOT achieves an accuracy gain of 8.6% on GSM8K from 57.54% to 66.19%, while DeepSeek-R1-Distill-Llama-70B with SLOT achieves a SOTA accuracy of 68.69% on GPQA among 70B-level models. Our code is available at https://github.com/maple-research-lab/SLOT.
title SLOT: Sample-specific Language Model Optimization at Test-time
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.12392