Representing data in words: A context engineering approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Caut, Amandine M., Rouillard, Amy, Zenebe, Beimnet, Green, Matthias, Morthens, Ágúst Pálmason, Sumpter, David J. T.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910050975481856
author Caut, Amandine M.
Rouillard, Amy
Zenebe, Beimnet
Green, Matthias
Morthens, Ágúst Pálmason
Sumpter, David J. T.
author_facet Caut, Amandine M.
Rouillard, Amy
Zenebe, Beimnet
Green, Matthias
Morthens, Ágúst Pálmason
Sumpter, David J. T.
contents Large language models (LLMs) have demonstrated remarkable potential across a broad range of applications. However, producing reliable text that faithfully represents data remains a challenge. While prior work has shown that task-specific conditioning through in-context learning and knowledge augmentation can improve performance, LLMs continue to struggle with interpreting and reasoning about numerical data. To address this, we introduce wordalisations, a methodology for generating stylistically natural narratives from data. Much like how visualisations display numerical data in a way that is easy to digest, wordalisations abstract data insights into descriptive texts. To illustrate the method's versatility, we apply it to three application areas: scouting football players, personality tests, and international survey data. Due to the absence of standardized benchmarks for this specific task, we conduct LLM-as-a-judge and human-as-a-judge evaluations to assess accuracy across the three applications. We found that wordalisation produces engaging texts that accurately represent the data. We further describe best practice methods for open and transparent development of communication about data.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Representing data in words: A context engineering approach
Caut, Amandine M.
Rouillard, Amy
Zenebe, Beimnet
Green, Matthias
Morthens, Ágúst Pálmason
Sumpter, David J. T.
Human-Computer Interaction
Computation and Language
Large language models (LLMs) have demonstrated remarkable potential across a broad range of applications. However, producing reliable text that faithfully represents data remains a challenge. While prior work has shown that task-specific conditioning through in-context learning and knowledge augmentation can improve performance, LLMs continue to struggle with interpreting and reasoning about numerical data. To address this, we introduce wordalisations, a methodology for generating stylistically natural narratives from data. Much like how visualisations display numerical data in a way that is easy to digest, wordalisations abstract data insights into descriptive texts. To illustrate the method's versatility, we apply it to three application areas: scouting football players, personality tests, and international survey data. Due to the absence of standardized benchmarks for this specific task, we conduct LLM-as-a-judge and human-as-a-judge evaluations to assess accuracy across the three applications. We found that wordalisation produces engaging texts that accurately represent the data. We further describe best practice methods for open and transparent development of communication about data.
title Representing data in words: A context engineering approach
topic Human-Computer Interaction
Computation and Language
url https://arxiv.org/abs/2503.15509