SIEVE: Sample-Efficient Parametric Learning from Natural Language

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Asawa, Parth, Dimakis, Alexandros G., Zaharia, Matei
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917382009651200
author Asawa, Parth
Dimakis, Alexandros G.
Zaharia, Matei
author_facet Asawa, Parth
Dimakis, Alexandros G.
Zaharia, Matei
contents Natural language context-such as instructions, knowledge, or feedback-contains rich signal for adapting language models. While in-context learning provides adaptation via the prompt, parametric learning persists into model weights and can improve performance further, though is data hungry and heavily relies on either high-quality traces or automated verifiers. We propose SIEVE, a method for sample-efficient parametric learning from natural language context that requires as few as three query examples. SIEVE uses a novel synthetic data generation pipeline, SIEVE-GEN, that leverages the insight that context is decomposable. Decomposing context allows us to generate higher quality rollouts by pairing synthetic queries with only the applicable context rather than the entirety, then using context distillation to internalize context into the model. We evaluate in reasoning settings where context is necessary, including custom domains and the RuleArena and Machine Translation from One Book tasks. Our results show that SIEVE outperforms prior context distillation methods using just three query examples, demonstrating how to achieve sample-efficient parametric learning from natural language.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02339
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SIEVE: Sample-Efficient Parametric Learning from Natural Language
Asawa, Parth
Dimakis, Alexandros G.
Zaharia, Matei
Machine Learning
Computation and Language
Natural language context-such as instructions, knowledge, or feedback-contains rich signal for adapting language models. While in-context learning provides adaptation via the prompt, parametric learning persists into model weights and can improve performance further, though is data hungry and heavily relies on either high-quality traces or automated verifiers. We propose SIEVE, a method for sample-efficient parametric learning from natural language context that requires as few as three query examples. SIEVE uses a novel synthetic data generation pipeline, SIEVE-GEN, that leverages the insight that context is decomposable. Decomposing context allows us to generate higher quality rollouts by pairing synthetic queries with only the applicable context rather than the entirety, then using context distillation to internalize context into the model. We evaluate in reasoning settings where context is necessary, including custom domains and the RuleArena and Machine Translation from One Book tasks. Our results show that SIEVE outperforms prior context distillation methods using just three query examples, demonstrating how to achieve sample-efficient parametric learning from natural language.
title SIEVE: Sample-Efficient Parametric Learning from Natural Language
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2604.02339