Language Model Inversion through End-to-End Differentiation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Denamganaï, Kevin Yandoka, Subr, Kartic
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914323203358720
author Denamganaï, Kevin Yandoka
Subr, Kartic
author_facet Denamganaï, Kevin Yandoka
Subr, Kartic
contents Despite emerging research on Language Models (LM), few approaches analyse the invertibility of LMs. That is, given a LM and a desirable target output sequence of tokens, determining what input prompts would yield the target output remains an open problem. We formulate this problem as a classical gradient-based optimisation. First, we propose a simple algorithm to achieve end-to-end differentiability of a given (frozen) LM and then find optimised prompts via gradient descent. Our central insight is to view LMs as functions operating on sequences of distributions over tokens (rather than the traditional view as functions on sequences of tokens). Our experiments and ablations demonstrate that our DLM-powered inversion can reliably and efficiently optimise prompts of lengths $10$ and $80$ for targets of length $20$, for several white-box LMs (out-of-the-box).
format Preprint
id arxiv_https___arxiv_org_abs_2602_11044
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Language Model Inversion through End-to-End Differentiation
Denamganaï, Kevin Yandoka
Subr, Kartic
Computation and Language
Artificial Intelligence
Despite emerging research on Language Models (LM), few approaches analyse the invertibility of LMs. That is, given a LM and a desirable target output sequence of tokens, determining what input prompts would yield the target output remains an open problem. We formulate this problem as a classical gradient-based optimisation. First, we propose a simple algorithm to achieve end-to-end differentiability of a given (frozen) LM and then find optimised prompts via gradient descent. Our central insight is to view LMs as functions operating on sequences of distributions over tokens (rather than the traditional view as functions on sequences of tokens). Our experiments and ablations demonstrate that our DLM-powered inversion can reliably and efficiently optimise prompts of lengths $10$ and $80$ for targets of length $20$, for several white-box LMs (out-of-the-box).
title Language Model Inversion through End-to-End Differentiation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.11044