Leveraging semantic similarity for experimentation with AI-generated treatments

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Lei, Arbour, David, Addanki, Raghavendra, Sinha, Ritwik, Feller, Avi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917039170387968
author Shi, Lei
Arbour, David
Addanki, Raghavendra
Sinha, Ritwik
Feller, Avi
author_facet Shi, Lei
Arbour, David
Addanki, Raghavendra
Sinha, Ritwik
Feller, Avi
contents Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main methodological challenge in this setting is representing these high-dimensional treatments without losing their semantic meaning or rendering analysis intractable. Here, we address this problem by focusing on learning low-dimensional representations that capture the underlying structure of such treatments. These representations enable downstream applications such as guiding generative models to produce meaningful treatment variants and facilitating adaptive assignment in online experiments. We propose double kernel representation learning, which models the causal effect through the inner product of kernel-based representations of treatments and user covariates. We develop an alternating-minimization algorithm that learns these representations efficiently from data and provides convergence guarantees under a low-rank factor model. As an application of this framework, we introduce an adaptive design strategy for online experimentation and demonstrate the method's effectiveness through numerical experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging semantic similarity for experimentation with AI-generated treatments
Shi, Lei
Arbour, David
Addanki, Raghavendra
Sinha, Ritwik
Feller, Avi
Methodology
Machine Learning
62K86, 65F55
Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main methodological challenge in this setting is representing these high-dimensional treatments without losing their semantic meaning or rendering analysis intractable. Here, we address this problem by focusing on learning low-dimensional representations that capture the underlying structure of such treatments. These representations enable downstream applications such as guiding generative models to produce meaningful treatment variants and facilitating adaptive assignment in online experiments. We propose double kernel representation learning, which models the causal effect through the inner product of kernel-based representations of treatments and user covariates. We develop an alternating-minimization algorithm that learns these representations efficiently from data and provides convergence guarantees under a low-rank factor model. As an application of this framework, we introduce an adaptive design strategy for online experimentation and demonstrate the method's effectiveness through numerical experiments.
title Leveraging semantic similarity for experimentation with AI-generated treatments
topic Methodology
Machine Learning
62K86, 65F55
url https://arxiv.org/abs/2510.21119