Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pal, Sayantan, Das, Souvik, Srihari, Rohini K.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916525086081024
author Pal, Sayantan
Das, Souvik
Srihari, Rohini K.
author_facet Pal, Sayantan
Das, Souvik
Srihari, Rohini K.
contents Large Language Models (LLMs) have significantly improved personalized conversational capabilities. However, existing datasets like Persona Chat, Synthetic Persona Chat, and Blended Skill Talk rely on static, predefined personas. This approach often results in dialogues that fail to capture human personalities' fluid and evolving nature. To overcome these limitations, we introduce a novel dataset with around 400,000 dialogues and a framework for generating personalized conversations using long-form journal entries from Reddit. Our approach clusters journal entries for each author and filters them by selecting the most representative cluster, ensuring that the retained entries best reflect the author's personality. We further refine the data by capturing the Big Five personality traits --openness, conscientiousness, extraversion, agreeableness, and neuroticism --ensuring that dialogues authentically reflect an individual's personality. Using Llama 3 70B, we generate high-quality, personality-rich dialogues grounded in these journal entries. Fine-tuning models on this dataset leads to an 11% improvement in capturing personality traits on average, outperforming existing approaches in generating more coherent and personality-driven dialogues.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11250
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations
Pal, Sayantan
Das, Souvik
Srihari, Rohini K.
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have significantly improved personalized conversational capabilities. However, existing datasets like Persona Chat, Synthetic Persona Chat, and Blended Skill Talk rely on static, predefined personas. This approach often results in dialogues that fail to capture human personalities' fluid and evolving nature. To overcome these limitations, we introduce a novel dataset with around 400,000 dialogues and a framework for generating personalized conversations using long-form journal entries from Reddit. Our approach clusters journal entries for each author and filters them by selecting the most representative cluster, ensuring that the retained entries best reflect the author's personality. We further refine the data by capturing the Big Five personality traits --openness, conscientiousness, extraversion, agreeableness, and neuroticism --ensuring that dialogues authentically reflect an individual's personality. Using Llama 3 70B, we generate high-quality, personality-rich dialogues grounded in these journal entries. Fine-tuning models on this dataset leads to an 11% improvement in capturing personality traits on average, outperforming existing approaches in generating more coherent and personality-driven dialogues.
title Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.11250