Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
Fuente:
arXiv
Saved in:
| Main Authors: | Truong, Kimberly Le, Fogliato, Riccardo, Heidari, Hoda, Wu, Zhiwei Steven |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation
by: Hu, Zhe, et al.
Published: (2024)
by: Hu, Zhe, et al.
Published: (2024)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026)
by: Le, Van-Truong
Published: (2026)
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
by: Bertran, Martin, et al.
Published: (2026)
by: Bertran, Martin, et al.
Published: (2026)
PersonaGym: Evaluating Persona Agents and LLMs
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
by: Liu, Shenyang, et al.
Published: (2025)
by: Liu, Shenyang, et al.
Published: (2025)
Automatic Legal Writing Evaluation of LLMs
by: Pires, Ramon, et al.
Published: (2025)
by: Pires, Ramon, et al.
Published: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
Distilling Text Style Transfer With Self-Explanation From LLMs
by: Zhang, Chiyu, et al.
Published: (2024)
by: Zhang, Chiyu, et al.
Published: (2024)
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
by: Ying, Shuangshuang, et al.
Published: (2025)
by: Ying, Shuangshuang, et al.
Published: (2025)
WritingBench: A Comprehensive Benchmark for Generative Writing
by: Wu, Yuning, et al.
Published: (2025)
by: Wu, Yuning, et al.
Published: (2025)
Not All Personas Are Worth It: Culture-Reflective Persona Data Augmentation
by: Han, Ji-Eun, et al.
Published: (2025)
by: Han, Ji-Eun, et al.
Published: (2025)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
Localizing Persona Representations in LLMs
by: Cintas, Celia, et al.
Published: (2025)
by: Cintas, Celia, et al.
Published: (2025)
Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
by: Wang, Zhengxiang, et al.
Published: (2025)
by: Wang, Zhengxiang, et al.
Published: (2025)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
by: Luo, Zhimeng, et al.
Published: (2025)
by: Luo, Zhimeng, et al.
Published: (2025)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
Steering at the Source: Style Modulation Heads for Robust Persona Control
by: Izawa, Yoshihiro, et al.
Published: (2026)
by: Izawa, Yoshihiro, et al.
Published: (2026)
Cascaded Self-Evaluation Augmented Training for Lightweight Multimodal LLMs
by: Lv, Zheqi, et al.
Published: (2025)
by: Lv, Zheqi, et al.
Published: (2025)
DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona
by: Choi, Janghyeok, et al.
Published: (2026)
by: Choi, Janghyeok, et al.
Published: (2026)
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
by: Zheng, Huaixiu Steven, et al.
Published: (2024)
by: Zheng, Huaixiu Steven, et al.
Published: (2024)
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
by: Fein, Daniel, et al.
Published: (2025)
by: Fein, Daniel, et al.
Published: (2025)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
by: Li, Zizhen, et al.
Published: (2025)
by: Li, Zizhen, et al.
Published: (2025)
PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
by: Pang, Tsz Fung, et al.
Published: (2025)
by: Pang, Tsz Fung, et al.
Published: (2025)
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
by: Lee, Huije, et al.
Published: (2026)
by: Lee, Huije, et al.
Published: (2026)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
by: Liu, Junlin, et al.
Published: (2026)
by: Liu, Junlin, et al.
Published: (2026)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
by: Wu, Weihao, et al.
Published: (2025)
by: Wu, Weihao, et al.
Published: (2025)
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
by: Zhu, Xiaoyuan, et al.
Published: (2026)
by: Zhu, Xiaoyuan, et al.
Published: (2026)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
by: Hua, Tianyu, et al.
Published: (2025)
by: Hua, Tianyu, et al.
Published: (2025)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
by: Liu, Wanlong, et al.
Published: (2024)
by: Liu, Wanlong, et al.
Published: (2024)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
by: Omar, Reham, et al.
Published: (2025)
by: Omar, Reham, et al.
Published: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
PARAN: Persona-Augmented Review ANswering system on Food Delivery Review Dataset
by: Park, Moonsoo, et al.
Published: (2025)
by: Park, Moonsoo, et al.
Published: (2025)
Vendi-RAG: Adaptively Trading-Off Diversity And Quality Significantly Improves Retrieval Augmented Generation With LLMs
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
Similar Items
-
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025) -
Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation
by: Hu, Zhe, et al.
Published: (2024) -
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026) -
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
by: Bertran, Martin, et al.
Published: (2026) -
PersonaGym: Evaluating Persona Agents and LLMs
by: Samuel, Vinay, et al.
Published: (2024)