StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shenyang, Gao, Yang, Zhai, Shaoyan, Wang, Liqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913780032602112
author Liu, Shenyang
Gao, Yang
Zhai, Shaoyan
Wang, Liqiang
author_facet Liu, Shenyang
Gao, Yang
Zhai, Shaoyan
Wang, Liqiang
contents Prompt Recovery, reconstructing prompts from the outputs of large language models (LLMs), has grown in importance as LLMs become ubiquitous. Most users access LLMs through APIs without internal model weights, relying only on outputs and logits, which complicates recovery. This paper explores a unique prompt recovery task focused on reconstructing prompts for style transfer and rephrasing, rather than typical question-answering. We introduce a dataset created with LLM assistance, ensuring quality through multiple techniques, and test methods like zero-shot, few-shot, jailbreak, chain-of-thought, fine-tuning, and a novel canonical-prompt fallback for poor-performing cases. Our results show that one-shot and fine-tuning yield the best outcomes but highlight flaws in traditional sentence similarity metrics for evaluating prompt recovery. Contributions include (1) a benchmark dataset, (2) comprehensive experiments on prompt recovery strategies, and (3) identification of limitations in current evaluation metrics, all of which advance general prompt recovery research, where the structure of the input prompt is unrestricted.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
Liu, Shenyang
Gao, Yang
Zhai, Shaoyan
Wang, Liqiang
Computation and Language
Artificial Intelligence
Machine Learning
Prompt Recovery, reconstructing prompts from the outputs of large language models (LLMs), has grown in importance as LLMs become ubiquitous. Most users access LLMs through APIs without internal model weights, relying only on outputs and logits, which complicates recovery. This paper explores a unique prompt recovery task focused on reconstructing prompts for style transfer and rephrasing, rather than typical question-answering. We introduce a dataset created with LLM assistance, ensuring quality through multiple techniques, and test methods like zero-shot, few-shot, jailbreak, chain-of-thought, fine-tuning, and a novel canonical-prompt fallback for poor-performing cases. Our results show that one-shot and fine-tuning yield the best outcomes but highlight flaws in traditional sentence similarity metrics for evaluating prompt recovery. Contributions include (1) a benchmark dataset, (2) comprehensive experiments on prompt recovery strategies, and (3) identification of limitations in current evaluation metrics, all of which advance general prompt recovery research, where the structure of the input prompt is unrestricted.
title StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.04373