Transmuting prompts into weights

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mazzawi, Hanna, Dherin, Benoit, Munn, Michael, Wunder, Michael, Gonzalvo, Javier
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910012565094400
author Mazzawi, Hanna
Dherin, Benoit
Munn, Michael
Wunder, Michael
Gonzalvo, Javier
author_facet Mazzawi, Hanna
Dherin, Benoit
Munn, Michael
Wunder, Michael
Gonzalvo, Javier
contents A growing body of research has demonstrated that the behavior of large language models can be effectively controlled at inference time by directly modifying their internal states, either through vector additions to their activations or through updates to their weight matrices. These techniques, while powerful, are often guided by empirical heuristics, such as deriving steering vectors from the average activations of contrastive prompts. This work provides a theoretical foundation for these interventions, explaining how they emerge from the fundamental computations of the transformer architecture. Building on the recent finding that a prompt's influence can be mathematically mapped to token-dependent implicit weight updates (Dherin et. al, 2025), we derive a principled method for condensing this information into token-independent thought vectors and thought matrices. These constructs provide a theoretical explanation for existing vector-and-matrix-based model editing techniques and offer a direct, computationally-grounded method for transmuting textual input into reusable weight updates.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08734
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transmuting prompts into weights
Mazzawi, Hanna
Dherin, Benoit
Munn, Michael
Wunder, Michael
Gonzalvo, Javier
Machine Learning
A growing body of research has demonstrated that the behavior of large language models can be effectively controlled at inference time by directly modifying their internal states, either through vector additions to their activations or through updates to their weight matrices. These techniques, while powerful, are often guided by empirical heuristics, such as deriving steering vectors from the average activations of contrastive prompts. This work provides a theoretical foundation for these interventions, explaining how they emerge from the fundamental computations of the transformer architecture. Building on the recent finding that a prompt's influence can be mathematically mapped to token-dependent implicit weight updates (Dherin et. al, 2025), we derive a principled method for condensing this information into token-independent thought vectors and thought matrices. These constructs provide a theoretical explanation for existing vector-and-matrix-based model editing techniques and offer a direct, computationally-grounded method for transmuting textual input into reusable weight updates.
title Transmuting prompts into weights
topic Machine Learning
url https://arxiv.org/abs/2510.08734