IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kantharuban, Anjali, Srivastava, Aarohi, Faisal, Fahim, Ahia, Orevaoghene, Anastasopoulos, Antonios, Chiang, David, Tsvetkov, Yulia, Neubig, Graham
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910105826492416
author Kantharuban, Anjali
Srivastava, Aarohi
Faisal, Fahim
Ahia, Orevaoghene
Anastasopoulos, Antonios
Chiang, David
Tsvetkov, Yulia
Neubig, Graham
author_facet Kantharuban, Anjali
Srivastava, Aarohi
Faisal, Fahim
Ahia, Orevaoghene
Anastasopoulos, Antonios
Chiang, David
Tsvetkov, Yulia
Neubig, Graham
contents Existing sentence representations primarily encode what a sentence says, rather than how it is expressed, even though the latter is important for many applications. In contrast, we develop sentence representations that capture style and dialect, decoupled from semantic content. We call this the task of idiolectal representation learning. We introduce IDIOLEX, a framework for training models that combines supervision from a sentence's provenance with linguistic features of a sentence's content, to learn a continuous representation of each sentence's style and dialect. We evaluate the approach on dialects of both Arabic and Spanish. The learned representations capture meaningful variation and transfer across domains for analysis and classification. We further explore the use of these representations as training objectives for stylistically aligning language models. Our results suggest that jointly modeling individual and community-level variation provides a useful perspective for studying idiolect and supports downstream applications requiring sensitivity to stylistic differences, such as developing diverse and accessible LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04704
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation
Kantharuban, Anjali
Srivastava, Aarohi
Faisal, Fahim
Ahia, Orevaoghene
Anastasopoulos, Antonios
Chiang, David
Tsvetkov, Yulia
Neubig, Graham
Computation and Language
Existing sentence representations primarily encode what a sentence says, rather than how it is expressed, even though the latter is important for many applications. In contrast, we develop sentence representations that capture style and dialect, decoupled from semantic content. We call this the task of idiolectal representation learning. We introduce IDIOLEX, a framework for training models that combines supervision from a sentence's provenance with linguistic features of a sentence's content, to learn a continuous representation of each sentence's style and dialect. We evaluate the approach on dialects of both Arabic and Spanish. The learned representations capture meaningful variation and transfer across domains for analysis and classification. We further explore the use of these representations as training objectives for stylistically aligning language models. Our results suggest that jointly modeling individual and community-level variation provides a useful perspective for studying idiolect and supports downstream applications requiring sensitivity to stylistic differences, such as developing diverse and accessible LLMs.
title IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation
topic Computation and Language
url https://arxiv.org/abs/2604.04704