Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kogan, David, Schumacher, Max, Nguyen, Sam, Suzuki, Masanori, Smith, Melissa, Bellows, Chloe Sophia, Bernstein, Jared
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909651323322368
author Kogan, David
Schumacher, Max
Nguyen, Sam
Suzuki, Masanori
Smith, Melissa
Bellows, Chloe Sophia
Bernstein, Jared
author_facet Kogan, David
Schumacher, Max
Nguyen, Sam
Suzuki, Masanori
Smith, Melissa
Bellows, Chloe Sophia
Bernstein, Jared
contents There is an unmet need to evaluate the language difficulty of short, conversational passages of text, particularly for training and filtering Large Language Models (LLMs). We introduce Ace-CEFR, a dataset of English conversational text passages expert-annotated with their corresponding level of text difficulty. We experiment with several models on Ace-CEFR, including Transformer-based models and LLMs. We show that models trained on Ace-CEFR can measure text difficulty more accurately than human experts and have latency appropriate to production environments. Finally, we release the Ace-CEFR dataset to the public for research and development.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications
Kogan, David
Schumacher, Max
Nguyen, Sam
Suzuki, Masanori
Smith, Melissa
Bellows, Chloe Sophia
Bernstein, Jared
Computation and Language
Artificial Intelligence
There is an unmet need to evaluate the language difficulty of short, conversational passages of text, particularly for training and filtering Large Language Models (LLMs). We introduce Ace-CEFR, a dataset of English conversational text passages expert-annotated with their corresponding level of text difficulty. We experiment with several models on Ace-CEFR, including Transformer-based models and LLMs. We show that models trained on Ace-CEFR can measure text difficulty more accurately than human experts and have latency appropriate to production environments. Finally, we release the Ace-CEFR dataset to the public for research and development.
title Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.14046