Subliminal Effects in Your Data: A General Mechanism via Log-Linearity

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Aden-Ali, Ishaq, Golowich, Noah, Liu, Allen, Shetty, Abhishek, Moitra, Ankur, Haghtalab, Nika
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908813580304384
author Aden-Ali, Ishaq
Golowich, Noah
Liu, Allen
Shetty, Abhishek
Moitra, Ankur
Haghtalab, Nika
author_facet Aden-Ali, Ishaq
Golowich, Noah
Liu, Allen
Shetty, Abhishek
Moitra, Ankur
Haghtalab, Nika
contents Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understand the effects of datasets on the model's properties. This is exacerbated by recent experiments that show datasets can transmit signals that are not directly observable from individual datapoints, posing a conceptual challenge for dataset-centric understandings of LLM training and suggesting a missing fundamental account of such phenomena. Towards understanding such effects, inspired by recent work on the linear structure of LLMs, we uncover a general mechanism through which hidden subtexts can arise in generic datasets. We introduce Logit-Linear-Selection (LLS), a method that prescribes how to select subsets of a generic preference dataset to elicit a wide range of hidden effects. We apply LLS to discover subsets of real-world datasets so that models trained on them exhibit behaviors ranging from having specific preferences, to responding to prompts in a different language not present in the dataset, to taking on a different persona. Crucially, the effect persists for the selected subset, across models with varying architectures, supporting its generality and universality.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04863
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
Aden-Ali, Ishaq
Golowich, Noah
Liu, Allen
Shetty, Abhishek
Moitra, Ankur
Haghtalab, Nika
Machine Learning
Artificial Intelligence
Computation and Language
Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understand the effects of datasets on the model's properties. This is exacerbated by recent experiments that show datasets can transmit signals that are not directly observable from individual datapoints, posing a conceptual challenge for dataset-centric understandings of LLM training and suggesting a missing fundamental account of such phenomena. Towards understanding such effects, inspired by recent work on the linear structure of LLMs, we uncover a general mechanism through which hidden subtexts can arise in generic datasets. We introduce Logit-Linear-Selection (LLS), a method that prescribes how to select subsets of a generic preference dataset to elicit a wide range of hidden effects. We apply LLS to discover subsets of real-world datasets so that models trained on them exhibit behaviors ranging from having specific preferences, to responding to prompts in a different language not present in the dataset, to taking on a different persona. Crucially, the effect persists for the selected subset, across models with varying architectures, supporting its generality and universality.
title Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.04863