PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cisar, Caitlin, Sheffield, Emily, Drake, Joshua, Harrell, Alden, Chidambaram, Subramanian, Nangia, Nikita, Arannil, Vinayak, Williams, Alex
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908546489122816
author Cisar, Caitlin
Sheffield, Emily
Drake, Joshua
Harrell, Alden
Chidambaram, Subramanian
Nangia, Nikita
Arannil, Vinayak
Williams, Alex
author_facet Cisar, Caitlin
Sheffield, Emily
Drake, Joshua
Harrell, Alden
Chidambaram, Subramanian
Nangia, Nikita
Arannil, Vinayak
Williams, Alex
contents Generative AI applications commonly leverage user personas as a steering mechanism for synthetic data generation, but reliance on natural language representations forces models to make unintended inferences about which attributes to emphasize, limiting precise control over outputs. We introduce PILOT (Psychological and Linguistic Output Targeting), a two-phase framework for steering large language models with structured psycholinguistic profiles. In Phase 1, PILOT translates natural language persona descriptions into multidimensional profiles with normalized scores across linguistic and psychological dimensions. In Phase 2, these profiles guide generation along measurable axes of variation. We evaluate PILOT across three state-of-the-art LLMs (Mistral Large 2, Deepseek-R1, LLaMA 3.3 70B) using 25 synthetic personas under three conditions: Natural-language Persona Steering (NPS), Schema-Based Steering (SBS), and Hybrid Persona-Schema Steering (HPS). Results demonstrate that schema-based approaches significantly reduce artificial-sounding persona repetition while improving output coherence, with silhouette scores increasing from 0.098 to 0.237 and topic purity from 0.773 to 0.957. Our analysis reveals a fundamental trade-off: SBS produces more concise outputs with higher topical consistency, while NPS offers greater lexical diversity but reduced predictability. HPS achieves a balance between these extremes, maintaining output variety while preserving structural consistency. Expert linguistic evaluation confirms that PILOT maintains high response quality across all conditions, with no statistically significant differences between steering approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15447
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting
Cisar, Caitlin
Sheffield, Emily
Drake, Joshua
Harrell, Alden
Chidambaram, Subramanian
Nangia, Nikita
Arannil, Vinayak
Williams, Alex
Computation and Language
Artificial Intelligence
Generative AI applications commonly leverage user personas as a steering mechanism for synthetic data generation, but reliance on natural language representations forces models to make unintended inferences about which attributes to emphasize, limiting precise control over outputs. We introduce PILOT (Psychological and Linguistic Output Targeting), a two-phase framework for steering large language models with structured psycholinguistic profiles. In Phase 1, PILOT translates natural language persona descriptions into multidimensional profiles with normalized scores across linguistic and psychological dimensions. In Phase 2, these profiles guide generation along measurable axes of variation. We evaluate PILOT across three state-of-the-art LLMs (Mistral Large 2, Deepseek-R1, LLaMA 3.3 70B) using 25 synthetic personas under three conditions: Natural-language Persona Steering (NPS), Schema-Based Steering (SBS), and Hybrid Persona-Schema Steering (HPS). Results demonstrate that schema-based approaches significantly reduce artificial-sounding persona repetition while improving output coherence, with silhouette scores increasing from 0.098 to 0.237 and topic purity from 0.773 to 0.957. Our analysis reveals a fundamental trade-off: SBS produces more concise outputs with higher topical consistency, while NPS offers greater lexical diversity but reduced predictability. HPS achieves a balance between these extremes, maintaining output variety while preserving structural consistency. Expert linguistic evaluation confirms that PILOT maintains high response quality across all conditions, with no statistically significant differences between steering approaches.
title PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.15447